
OpenAI Codex User Alleges Unauthorized Agent Charges
Published by AINave Editorial
A Codex user says a request for a UX/UI review led to hundreds of child tasks and nearly $80,000 in invoices. The account raises a traceability problem: local task records and a reconstructed invoice history do not, by themselves, establish what happened on OpenAI’s servers or what caused the charges.
What the user says the task did
Lorenzo Massaro says he started a Codex task in VS Code on July 10, 2026, asking for UX/UI validation of a product module. He says the GPT-5.5 task, set to Medium reasoning, created 826 distinct child-task records, listed as GPT-5.6 Sol with Ultra reasoning. The child-task titles, according to his post, ranged beyond UI review into topics including backend infrastructure, OAuth, audits and release work.
Those details are Massaro’s account on Hacker News, not an independently verified incident report. The supplied evidence contains no OpenAI confirmation that the tasks were unauthorized or that Codex behaved as he describes.
The invoice total is not a token-count calculation
Massaro says his reconstructed billing history contains 162 paid invoices totaling $79,664.88. The post’s headline rounds the allegation to $78,000, so the exact figure should be understood as the total he says he reconstructed, not a confirmed OpenAI billing statement.
His local task data does not settle the bill. Massaro explicitly says local token counters are not OpenAI’s authoritative billing ledger and cannot simply be multiplied by an API price. He also says roughly 2,550 non-archived legacy threads retained metadata but lacked corresponding raw rollout files on his machine. That gap matters: metadata may show that records existed, but without the detailed execution history, it is harder to reconstruct what generated them.
A version correlation, not a demonstrated cause
Massaro compares the task records associated with two Codex builds. His post reports more local token-counter volume per child task under 0.144.0-alpha.4, but the comparison is based on his recovered local data, not an authoritative usage report.
| Build recorded in local data | Child tasks | Local token counters |
|---|---|---|
| 0.144.0-alpha.4 | 584 | About 154.36 billion |
| 0.144.2 | 242 | About 7.51 billion |
He calculates an approximately 8.5-fold difference in average local token volume per child and says 103 of 104 high-volume tasks were recorded under the alpha build. That is a reported correlation, not proof that the build caused the activity or that the counters map to billed tokens.
The missing piece is a server-side account
Massaro says he opened OpenAI Support case #15189838 and asked for a server-side reconstruction. In his account, the response said credits had been consumed but did not provide further details. The supplied material includes no OpenAI explanation or authoritative breakdown of the alleged charges.
The practical distinction is between a local trace of task records and the service-side evidence needed to connect those records to billed usage. Until that connection is made, the version comparison may point to where an investigation should look, but it cannot establish either the cause or the final bill.






















