A failure worth fixing once.
Three teams hit the same broken integration environment. 47 sessions worked around it one at a time, each burning time and tokens on the same dead end.
docker compose up exited before tests began
Understand your bill
Every month the invoice says your coding agents worked hard. It cannot say on what, for whom, or whether any of it was worth it.
Quesma turns Claude Code, Codex, and Cursor sessions into a clear view of what the agents did, what it cost, and what you got back.
Fix the flaky checkout retry without changing the API$0.000Read(payments/retry.ts)$3.72Bash(pnpm test checkout --filter retry)$8.40Read(tmp/checkout-ci.log)$22.10Read(tmp/checkout-ci.log) ↻ repeated context$34.76Edit(payments/retry.ts)$45.18Bash(pnpm test checkout --filter retry)$58.9242 passed · checkout retry fixed$64.28Behind the bill
Quesma reads the sessions themselves and sorts thousands of them into one picture of the work: by team, by repository, by kind of work.

A team you didn’t expect found a use for the tools.
The default model got smarter, and pricier.
The work ran hot for three weeks because you asked it to.
The patterns
The patterns only appear when sessions are analyzed together.
Cost tools only find things to cut. Quesma also finds problems recurring across teams, and the skills, tools, and routines used in one part of the company but nowhere else.
Three teams hit the same broken integration environment. 47 sessions worked around it one at a time, each burning time and tokens on the same dead end.
docker compose up exited before tests began
Custom skill deploy-preview lived on one team for months. Quesma shows where it is used, and which teams have not met it yet.
One repository’s rules send every session down the same wrong turn. Rewriting one file fixes all the future ones.
“Always run the full e2e suite before committing”
31 sessions took the same wrong turn
An MCP server spread session by session until half the company relied on it. Better to know.
first seen Feb 12 · no decision on record
The evidence
Frontier-model training, independent audits, open-source infrastructure, and research cited by the labs moving the field forward.
Can coding agents find backdoors in compiled binaries using Ghidra and radare2?
Explore BinaryAuditWe built training environments for difficult, multi-hour agent tasks. The work required reading full trajectories: prompts, tool calls, failures, retries, and results.
We audited its 66.5% SWE-Bench Pro result, reviewing the environment, required changes, and final score.
Read the independent audit OPEN-SOURCE CONTRIBUTIONTERMINAL-BENCH 3.0Quesma contributed to Terminal-Bench 3.0, a benchmark for testing agents on practical terminal tasks.
View the contributors RESEARCH / INDUSTRY CITATIONBABA IS BENCHOpenAI cited our work while examining what makes agent evaluations succeed.
Read the citationRESEARCH NOTES, IN PUBLIC
Quantization is a lossy compression, and factual knowledge is incompressible. We test Qwen3.6 27B GGUF quantizations from Hugging Face (by Unsloth and Bartowski) and llama.cpp on the Incompressible Knowledge Probes (IKP) benchmark.

We evaluate July 2026 fresh releases Kimi K3, Claude Opus 5, Grok 4.5, Gemini 3.6 Flash, and DeepSeek V4 Flash 0731 on Baba Is Bench, an LLM agent benchmark based on the puzzle game Baba Is You, comparing pass rate, speed, and cost with Claude Fable 5 and GPT-5.6.
A warning hidden in the DeepSeek-V4 paper says retrying interrupted LLM requests is mathematically incorrect — it introduces length bias. I reproduced it on 100,000 poems.
DESIGN PARTNERS
Early access opens soon. The collector deploys in your environment, the record accumulates in your cloud, and early teams get every analysis first. Contact us to get in early.