
GLM improved its inference. Engineers still set the goal
Z.ai reports 3× serving throughput with help from a GLM coding agent. I examine what improved, what humans controlled, and what developers can reuse.
Every article tagged coding-agents, newest first.

Z.ai reports 3× serving throughput with help from a GLM coding agent. I examine what improved, what humans controlled, and what developers can reuse.

Codex CLI 0.155 adds experimental voice and safer worktree cleanup. I explain the fixes for lost prompts, restarted sessions and the next-day hotfix.

DeepSeek V4.1 Flash activates 8B parameters on input, but its full weights are enormous. I explain the API prices, coding scores and server requirements.

Fable 5.1 leads its predecessor on long agent tasks, but Opus 5 stays close at half the base token price. I checked the benchmarks and first reports.

Codex gives me more usage, cleaner frontend work, and a smoother app. Claude still finishes my large multi-agent features in hours instead of all day.

After five weeks of daily use, I think Opus 5 is much better than its online reputation. Clear plans and smaller tasks reveal a strong coding model.

Cursor's Agents Window can open first, but Editor Window remains supported. I map the commands and startup setting that put code back in front of you.

DeepSWE puts Luna Max 2.2 points behind Sol High at roughly one-sixth the attempt cost. I explain why the models can still feel far apart in repository work.

I built a four-layer context management system that routes coding agents to current facts. It works, but stale documentation is still the hard part.

I compared Sol, Terra, Opus 5 and Fable 5 across coding benchmarks. The winner changes with the task, effort setting, agent setup and budget.

Kimi K3 ties GPT-5.6 medium but takes 4.6 times as long. GLM-5.2 is cheap per token yet costly per task. I checked where both models still win.