Case study

One changed ID: how an AI pentest became full account takeover

By Sintropyc · Case study

A real engagement, run end-to-end by Sintropyc's autonomous AI penetration-testing agent. It started from a single enumerable object ID and finished at the private admin dashboard — and along the way it threw away its own first finding as a likely false positive. All identifying details are anonymized; the app was tested with the owner's authorization.

Run a scanIDOR testing →

The one-line version

An autonomous AI agent tested an e-commerce SaaS the way a real attacker would. It found it could read other members' private profiles by changing one integer, rejected that finding as a probable false positive, re-proved it against genuinely separate accounts, and then walked the app all the way to admin — reading the private revenue dashboard as a customer it had registered sixty seconds earlier. Every step is a proven effect, not a scanner's "possible."

Authorized · anonymized

Testing was authorized by the application owner and run read-only by default in an isolated sandbox, using neutered, reversible payloads. The target, its domain, real users and any secrets are anonymized or omitted — nothing here identifies a real user, credential or system.

Why this is the bug that keeps winning

Broken Object-Level Authorization — BOLA, the API-native cousin of the classic IDOR — is number one on the OWASP API Security Top 10. The bug is boring and everywhere: an endpoint like GET /api/members/{id} returns the object for whatever id you pass, without checking that you are allowed to see it. Change the number, read someone else's data.

Scanners are bad at it. A scanner sees GET /api/members/42 return 200 OK and shrugs — 200 is what it's supposed to return. Whether the record belongs to you or to a stranger is a question about identity and business context, not HTTP status codes. That's why this is the class human testers find most and automated tools miss most — and the class our agent proves most often.

The signal: one changed ID

The target was a storefront SaaS — a single-page app talking to a JSON API. The agent mapped the API and noticed two things that matter for authorization bugs: member IDs were sequential integers, and the member object carried a password-recovery hint — a private field that should never leave the owner's own session.

So it registered Account A and asked a simple question: as A, can I read member 1? And 10? Both returned full profiles, recovery-hint field included. On paper: textbook IDOR. A scanner, or an eager tester, would screenshot the 200 and write it up as HIGH.

Our agent didn't.

The false positive it threw away

Before recording anything, the agent applied the standard it holds every authorization claim to: a cross-account read is only a real break if the two identities are genuinely distinct owners with no legitimate shared access to a private resource. And it spotted a hole in its own evidence — on many apps, self-registered accounts silently land in the same default workspace. If A and the "victim" share a workspace, A reading that record isn't a broken boundary; it's authorized shared access. A real vulnerability and a normal feature can look byte-for-byte identical in the response.

The difference

So the agent downgraded and rejected its own claim, and set a harder task: prove it with owners that are unambiguously separate. This is the single biggest source of false positives in access-control testing.

It registered a second, independent account, confirmed the two were distinct owners with no shared membership, and re-ran the cross-read. It held. One account could read another owner's private profile with nothing but a changed integer in the URL. Now it was a finding — severity set to exactly what was shown.

From a read to the whole store

A read primitive is bad. What made this a takeover was the next question: if the server doesn't check who can read an object, does it check who can write one — and what they can write? It didn't.

The registration and profile-update endpoints accepted a client-supplied role field and persisted it as-is. The agent registered a fresh customer and set its own role to admin. To prove the escalation was real, it used the privilege: as a customer minutes old, it opened the private admin dashboard — total members and revenue — and got a 200 OK. Then it repeated the trick across users: one customer flipped a different customer's role. No ownership check, no role check, on the write path either. Every change was reversed with benign, restore-on-completion markers.

The doors that had no lock at all

The agent also tried the front doors with no key. Several /admin/* endpoints and a user-listing endpoint returned data to a completely unauthenticated request: store configuration, the revenue dashboard, and every member's email and full name. No token, no session — just a GET. It also found the classic token flaw: the server accepted JWTs with alg: none and honored forged claims. Where it couldn't observe a signing secret, it said so and marked that part inferred, not proven.

How to make sure your app doesn't do this

The fixes are unglamorous and they work:

  • Authorize every object access, server-side, default-deny. On every /{id} route, verify the caller may see or change that specific object.
  • Never bind client input to privileged fields. Allow-list writable fields; role, tier, owner are set by the server via admin-only paths.
  • Treat authorization as owner-vs-owner. Test with two genuinely separate accounts.
  • Fix the token layer. Reject alg: none, pin the algorithm, verify signatures, rotate leaked secrets.
  • Assume every endpoint is public until proven authorized — especially /admin and /internal.

The point isn't the bug — it's the proof

Every finding here was demonstrated by effect, not asserted from a signature. A returned foreign record. A role change that persisted and opened a door. An unauthenticated 200 on data that needed a session. And when the evidence didn't meet that bar, the agent said so and either re-proved it or capped the claim. Proof, not guess. Severity never higher than what was actually shown.

See it proven on your own app

One fixed price. A full test, every finding proven with a working exploit, the exact fix, and a retest after you ship.