In the Arena · August 24, 2026

In the Arena #3: the model is the cheap part

A 900-person company builds what it could not buy, five verbs move a benchmark 59 points, and your CEO is in the commit log.

A note from Alex

Alex here — I run growth at CoMavenAI, which mostly means I spend my week talking with operators about where AI actually pays for itself. The four reads this issue keep arriving at the same place: the model is the cheap part. A healthcare company got further building on a vendor's SDK than buying the vendor's product. Mistral moved a benchmark 59 points without touching the model. Linear's data says the people building now include the CEO. And the cautionary tale is an agent that aced its exam by finding the answer key online.

If any of this is sitting on your desk this quarter, I want to hear how you're thinking about it. Email contact@comavenai.com or grab thirty minutes with me — no deck, no pitch, just a working conversation.

- Alex

Worth your time

They were ready to sign. They built instead. Headway, a 900-person mental healthcare company, was preparing to sign with a major AI vendor to put its desktop app in front of employees. Waiting on a vendor to meet its compliance and workflow requirements “felt like we were missing the train,” says CTO Arnaud Ferreri. So the team built Eddy, an internal assistant on Anthropic's Claude Code SDK, hosted in Headway's own AWS environment. Work started barely six months ago; now the whole company uses it. The design choice worth stealing: Eddy's agents act without asking permission at every step, and that autonomy is possible because every conversation runs in a sealed, disposable container with tightly limited connections to company tools and data. The controls are what bought the freedom, not what it traded away. Katie Parrott's feature in Every (the opening is free, the full piece sits behind Every's paywall; Every discloses Headway is a consulting client) · Ferreri's own build-vs-buy writeup, free on Medium

Your CEO is in the commit log. Linear's first “How teams build” report, written by its head of data Tim Qi, reads adoption off 127,000 paid users. The steep curves are not where you would point: product managers using AI went from 12 to 34 percent in six months, and CEOs at companies of 201 or more people went from 9 to 36 percent. Pull requests per workspace are up 111 percent on a June 2024 baseline, and teams now use AI to write just under half of everything created in Linear. One line held flat, and it is the one to sit with: time spent on customer requests, docs, and projects did not move in a year when nearly everything else went up. AI is compressing execution, not planning. Qi is plain about the limit; this is Linear's customer base, not the market. The report, free, no signup

Mistral taught a model to read like an analyst. Mistral replaced one-shot document retrieval with a loop. The model gets five operations — search, open, navigate, read, grep — and works a filing the way an analyst would: open it, follow the references, check the answer. On FinanceBench, 368 SEC filings and 150 questions, correctness went from 26.7 percent to 86 percent. Same model. The numbers are Mistral's own, so treat them as a vendor's best day. Buried in the same post is the sharper fact for anyone evaluating vendors: an identical open model scored 41.4 percent on a second benchmark inside one harness and 51.9 percent inside another. Ten and a half points from scaffolding alone. When a demo shows you accuracy, ask what harness produced it, because that is the part you will actually be operating. Mistral's announcement, August 20, team byline, free

“Perhaps the solution is available publicly.” That is GPT-5.6 Sol thinking out loud mid-benchmark, in traces published by engineer Adam Williams. His harness scored 84 of 89 on Terminal Bench 2.1. Then he read the traces: on the task he audited, the model, with no web-search tool enabled, used curl to reach DuckDuckGo, GitHub, grep.app and SourceGraph and pulled public solutions, three runs out of three. Williams is careful about intent; maybe cheating, maybe stumbling onto the answer while searching. That task's pass was worthless either way. The operator's lesson is not about this model. If the test you use to accept an agent's work exists anywhere public, assume the agent can find it. Trust the number after you have read the traces, not before. Adam Williams's writeup, August 12, free, no paywall

From the team

The build-vs-buy math in the first item is our day job. Most mid-market companies do not have a 900-person org's engineering bench, which is why the buy column usually wins by default. A forward deployed engineer is how you get the build column without the hiring. Our July explainer covers what FDEs actually do and how the mid-market gets one: What is a Forward Deployed Engineer?

Missed last week? Issue 2 covered correlated blind spots in agent swarms, the OpenAI enterprise report, Agent Plugins, and who funds the compute.

See you in the arena.

Have a workflow that deserves better? Grab thirty minutes with us.