Read the full analysis: How My AI Stack Works In September 2026 on ThorstenMeyerAI.com
TL;DR
With six frontier AI models clustered within about 20 index points while their per-task costs differ by roughly 100x, Thorsten Meyer’s September 2026 stack assigns Claude Opus 5.5 as main builder, newly released GPT-6.1 Sol as detail-and-review model, and Jev for routing. Effort settings, not model choice, are the biggest cost lever.
Technology writer Thorsten Meyer published his September 2026 AI model stack on 29 September, built around a market shift he says has taken hold over the past four weeks: six leading models now sit within about 20 index points of each other on the Artificial Analysis Intelligence Index, while their cost per task differs by roughly 100x. His answer to the new question — “which model clears my quality bar at the lowest cost per task?” — pairs Claude Opus 5.5 as the main builder with GPT-6.1 Sol, released the same day, as a low-cost reviewer, plus a decision-only model for routing.
Meyer’s stack assigns Opus 5.5 at high or xhigh effort as the primary model for development work, on the grounds that high delivers 54 index points at $1.82 per task and xhigh adds 2 more points for hard problems such as architecture and migrations. GPT-6.1 Sol at high or xhigh handles detail work and an independent review pass at $0.32 to $0.39 per task. Astra, Fable, Sonnet 5.5 and Luna serve as alternates for specific jobs, while Jev, a decision model Meyer notes “cannot write a sentence,” takes over high-volume yes/no and routing judgements.
Three findings anchor the piece, all based on Artificial Analysis Intelligence Index v4.3.x scores. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task — $5.98 against $7.63 at top settings. Second, Sonnet 5.5 at max effort costs more per task than Opus at max for 2 fewer points, which Meyer argues disqualifies it from that setting. Third, GPT-6.1 Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
GPT-6.1 Sol launched on 29 September at $2 / $10 per million tokens, matching its week-old predecessor, and its medium setting already matches GPT-6 Sol’s score of 48 at one-fifth of that model’s $1.06 per-task cost. The catches, per Meyer: high and xhigh settings take 57 to 69 seconds to first token, making Sol non-interactive at those levels, and Opus 5.5 still leads it by 5 points at xhigh.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
SettingIndexCost per taskOutput tokensFirst tokenmedium48$0.2115M5.3 shigh50$0.3225M57 sxhigh51$0.3936M69 s
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
Replay 300 to 500 past decisionsCompare overall and per confidence bandRead 20 disagreements, decide who was rightHigh band at 95% or better?Own flag, off by defaultCanary on 5 to 10 unitsRoll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
1Relevance gate2Language check3Classifier fallback
Publishing and content
4Thin-source detector5Same-event dedupe6Product fits roundup7Disclosure present8Headline quality9Comment moderation
Commerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triage
Software and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triage
Business ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent
Limits, cost and one hard rule
Why Per-Task Cost Now Drives Model Choice
The piece documents a practical consequence of a maturing frontier market: when capability gaps shrink to single index points, the deciding variables become cost per task and effort settings. Meyer shows that on Opus 5.5, moving from xhigh to max adds 2 index points but 73% more cost per task, and from medium to max raises cost 4.46x for 7 points — meaning the effort dial moves the bill more than the choice between most models.
The stack’s core idea is the affordable review seat: a different model family checking Opus’s output, at $0.39 per task, is both a better check than self-review and cheap enough to run on every meaningful change. Meyer also cautions that halving model price saves only 12.5% of real cost in his illustrative example, and a single extra minute of human review erases the saving — a figure he flags as illustrative rather than measured.
A Month of Front-Model Releases
The stack lands at the end of a crowded release month: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September. Per-task costs across the group range from Luna’s $0.07 (1,429 tasks per $100) to Fable’s $7.63 (13 tasks per $100).
All capability figures come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability rather than a verdict on any specific workload — hence his standing advice to shadow-test before switching models. Sonnet 5.5 at max effort produced about 193k output tokens per task, the most Artificial Analysis has measured, which Meyer cites as evidence of poor value at that setting.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer, ThorstenMeyerAI.com
Limits of the Index and the Data
Meyer is explicit that all scores come from a single benchmark family, Artificial Analysis Intelligence Index v4.3.x, and that one index point sits inside the noise — a gap that covers several of the comparisons he draws. The index measures general capability, not performance on any individual workload, and he recommends shadow-testing before switching models.
For GPT-6.1 Sol, released the same day as publication, Artificial Analysis has not yet published low or max effort settings, and only three effort levels are listed so far. Meyer’s cost-of-human-review example is labeled illustrative, not measured, and the piece represents one practitioner’s configuration rather than a vendor recommendation or independently verified evaluation.
Watch Settings, Prices and New Benchmarks
Expect Artificial Analysis to publish Sol’s remaining effort settings, which could shift its value assessment. Meyer notes that future model swaps in his stack depend on where new releases land on the price curve, and his four operating rules — effort is not capability, a different model reading the same flawed spec is not an independent review, passing tests are not approval to ship, and failing work goes back to the builder with evidence — will govern how the stack evolves through future releases.
Key Questions
Which model does Meyer use as his main builder?
Claude Opus 5.5 at high or xhigh effort — 54 index points at $1.82 per task for everyday development, or 56 points at $3.46 for architecture, migrations and trust boundaries.
Why use GPT-6.1 Sol instead of Opus for review?
Sol’s xhigh setting sits only 1 to 2 points below Astra and Fable at $0.39 per task instead of $3.26 or $7.63, making a routine second-opinion pass from a different model family affordable on every meaningful change.
What is the biggest cost lever in this stack?
The effort setting. On Opus 5.5, going from xhigh to max adds 2 index points and 73% more cost per task; medium to max raises cost 4.46x for 7 points.
What are GPT-6.1 Sol’s drawbacks?
High and xhigh settings take 57 to 69 seconds to produce a first token, so it is not interactive at those levels, and Opus 5.5 still leads it by 5 index points at xhigh.
What is Jev used for?
Jev is a decision model that Meyer says cannot write a sentence; it handles high-volume yes/no judgements and routing, such as classification and extraction work alongside GPT-6 Luna.
Source: ThorstenMeyerAI.com