
OpenAI confirms an Agents API container billing bug
A Reddit user reported a $1,600 Agents API bill. OpenAI confirmed container overbilling. I separate the resolved incident from the cleanup work apps need.
When a lab ships a model: what it actually changes for people who build software.

A Reddit user reported a $1,600 Agents API bill. OpenAI confirmed container overbilling. I separate the resolved incident from the cleanup work apps need.

Z.ai reports 3× serving throughput with help from a GLM coding agent. I examine what improved, what humans controlled, and what developers can reuse.

GPT-6 Sol and Luna cut API prices while independent benchmarks show smaller gains. I compare the scores, ChatGPT access, and mixed first reports on Reddit.

Opus 5.5 gains ground in independent coding tests and cuts token prices. I compare the task costs, effort settings, and first Reddit reports on daily use.

Anthropic says Claude leads a quarter of its AI research work under human supervision. I look at what the figure measures and who still checks the results.

Codex CLI 0.155 adds experimental voice and safer worktree cleanup. I explain the fixes for lost prompts, restarted sessions and the next-day hotfix.

Reddit users suspect a hidden Fable upgrade. I traced the claims to their sources and checked what would actually establish a new Claude model release.

DeepSeek V4.1 Flash activates 8B parameters on input, but its full weights are enormous. I explain the API prices, coding scores and server requirements.

OpenAI now hosts the Codex agent loop for your application. I explain how the Agents API differs from the SDK, what persists and which costs remain.

OpenAI's GPT-6 Astra is metered at 2.5 times Sol's rate and feels more human in conversation. I checked the limits, the benchmarks, and the first reactions.

Fable 5.1 leads its predecessor on long agent tasks, but Opus 5 stays close at half the base token price. I checked the benchmarks and first reports.

OpenAI's five-hour gate can interrupt concentrated Codex work while weekly usage remains. It limits when Plus can work, not necessarily total usage.

Claude models will watermark their text through the words they pick, not hidden characters. I trace the reason to the EU AI Act, not the distillation war.

OpenAI confirmed Astra as its next major model, then slowed internal work over cyber risk. The GPT-6 launch date came from a withdrawn leak and a joke.

OpenAI reset Codex usage five times in ten days while Sol drained limits faster than expected. I think the free refills helped hide the real usage problem.

V4 Flash costs $0.14 in and $0.28 out per million tokens. I checked whether its 97 to 99 percent discount makes DeepSeek's benchmark losses worth it.

Opus 5 matches Fable 5 on benchmarks at half the token price, yet it can be painful to supervise. Here is where Fable still earns its place.