
OpenAI confirms an Agents API container billing bug
A Reddit user reported a $1,600 Agents API bill. OpenAI confirmed container overbilling. I separate the resolved incident from the cleanup work apps need.
Model releases, benchmarks, experiments and guides for developers who work with AI every day. Written from real use, with sources for every claim.

A Reddit user reported a $1,600 Agents API bill. OpenAI confirmed container overbilling. I separate the resolved incident from the cleanup work apps need.

Z.ai reports 3× serving throughput with help from a GLM coding agent. I examine what improved, what humans controlled, and what developers can reuse.

GPT-6 Sol and Luna cut API prices while independent benchmarks show smaller gains. I compare the scores, ChatGPT access, and mixed first reports on Reddit.

Opus 5.5 gains ground in independent coding tests and cuts token prices. I compare the task costs, effort settings, and first Reddit reports on daily use.

Anthropic says Claude leads a quarter of its AI research work under human supervision. I look at what the figure measures and who still checks the results.

Codex CLI 0.155 adds experimental voice and safer worktree cleanup. I explain the fixes for lost prompts, restarted sessions and the next-day hotfix.

Reddit users suspect a hidden Fable upgrade. I traced the claims to their sources and checked what would actually establish a new Claude model release.

Claude Code's promotion ended on 14 September. I explain the smaller weekly allowance, the 25% versus 16.7% calculation and what five-hour warnings mean.

DeepSeek V4.1 Flash activates 8B parameters on input, but its full weights are enormous. I explain the API prices, coding scores and server requirements.

OpenAI now hosts the Codex agent loop for your application. I explain how the Agents API differs from the SDK, what persists and which costs remain.

OpenAI's GPT-6 Astra is metered at 2.5 times Sol's rate and feels more human in conversation. I checked the limits, the benchmarks, and the first reactions.

Muse Spark 1.3 scored 62 on Artificial Analysis, above GPT-6 Astra, then fell to ninth in a rescore. I traced where the score came from and what it costs.