🔍 Read the full analysis: A Role-Based Guide To My September 2026 AI Stack on ThorstenMeyerAI.com
Turn your quiet moments into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
TL;DR
A September 29 guide assigns Claude Opus 5.5 to building and newly released GPT-6.1 Sol to detailed review, while cheaper models handle routine tasks. The author’s comparisons use Artificial Analysis Intelligence Index v4.3.x and task-cost estimates; they are one person’s workflow, not proof of performance on every workload.
Thorsten Meyer published a role-based guide to his AI stack on September 29, assigning Claude Opus 5.5 to software building and newly released GPT-6.1 Sol to detailed review and analysis. The guide argues that with several models scoring within about 20 points on the cited benchmark but differing widely in estimated task cost, users should choose by workload and price rather than leaderboard rank alone.
Meyer says he uses Opus 5.5 at high effort for features, APIs, multi-file work and refactors, and xhigh for harder architectural work. In the Artificial Analysis Intelligence Index v4.3.x figures he cites, Opus scores 54 at high effort for an estimated $1.82 per task, and 56 at xhigh for $3.46. Its maximum setting scores 58, but raises the estimated task cost to $5.98.
GPT-6.1 Sol, released September 29, is his choice for examining files and reviewing changes. The cited index lists it at 48 points and $0.21 per task at medium effort, 50 points and $0.32 at high, and 51 points and $0.39 at xhigh. Meyer says the higher settings take 57 to 69 seconds to produce a first token, making them less suited to interactive use.
The guide assigns Sonnet 5.5 and Luna to scoped subtasks and routine checks, while Astra or Fable serve as alternatives when tests favor them or Sol and Opus disagree. Meyer says the index is a general capability measure rather than a verdict on a particular workload, and recommends shadow-testing before switching systems.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Task Costs Shape the Stack
The guide’s main practical claim is that task cost and quality need to be weighed together. Meyer’s table places Opus 5.5 at the top score, while Sol and Luna have much lower estimated costs per task. For teams running repeated reviews, extraction or routing, the cost gap could affect how often they can use a model, but the figures are benchmark estimates and do not establish savings for every organization.
It also highlights effort settings as a major cost choice. In Meyer’s cited results, moving Opus from medium to max increases task cost from $1.34 to $5.98 while its score rises from 51 to 58. He argues that max effort is rarely worthwhile for his work. This makes the guide relevant to developers and teams deciding whether extra model effort delivers enough quality improvement to justify its price.
Meyer recommends a separate model family for review, using Sol to check Opus’s work. That can provide another perspective, but independent review is not guaranteed: the guide itself cautions that two models can share the same flawed requirements. It says passing tests alone should not be treated as approval to ship.
Benchmarks Behind the September Guide
The article is a dated account of one user’s workflow, published September 29, 2026. Its model scores and estimated task costs are attributed primarily to Artificial Analysis Intelligence Index v4.3.x. Meyer describes that index as a map of general capability and advises readers to test models on their own tasks before changing systems.
The cited comparison lists Opus 5.5 at 58 points and $5.98 per task at its top setting; Sonnet 5.5 at 56 and $7.60; Fable 5.1 at 53 and $7.63; GPT-6 Astra at 53 and $3.26; GPT-6.1 Sol at 51 and $0.39; and GPT-6 Luna at 37 and $0.07. These are index figures and estimated task costs, not a guarantee of comparable results in a reader’s software, documents or business process.
Meyer also compares token prices: Opus at $4 input and $20 output per million tokens, Sol at $2 and $10, and Luna at $0.10 and $0.50. He warns that token price alone does not determine the cost of completed work, since additional human review time can outweigh a model-price saving.
““The practical reading: Sol is not the model I ask to build. It is the model I can afford to run on everything.””
— Thorsten Meyer
Limits of the Model Comparisons
The guide does not establish whether its benchmark rankings transfer to other workloads. Meyer says readers should shadow-test models before switching. The source also says one index point is within the noise, and that Artificial Analysis had not published GPT-6.1 Sol’s low or max effort results at the time of writing.
The excerpt does not provide a full method for the estimated cost-per-task calculations or independent measurements of the author’s workflow. Its final example about human review cost is marked illustrative, not measured, and the source text cuts off before completing it. Actual end-to-end costs and quality remain workload-dependent; the guide supplies no results from a controlled comparison across organizations.
Test the Stack on Real Work
Meyer’s recommended next step for readers is to shadow-test candidate models on their own tasks before replacing an existing workflow. Teams can compare output quality, latency and total cost, including the time people spend checking results, against their own requirements.
The GPT-6.1 Sol comparison may change as additional effort-level results become available in the index. Until then, the guide’s role assignments remain the author’s September 29 snapshot, not a settled ranking for all uses.
Key Questions
What is the main development in the guide?
Thorsten Meyer published a role-based AI workflow on September 29, 2026, including GPT-6.1 Sol, released that day, as a model for detailed review and analysis.
Which models does Meyer use for building and review?
He assigns Opus 5.5 to primary building and GPT-6.1 Sol to file-level analysis and review, using different effort settings depending on task difficulty.
Are the cost figures guaranteed for other users?
No. They are estimated task costs in the cited Artificial Analysis index comparison. Meyer says the index does not decide which model fits a specific workload and recommends shadow-testing.
Why does Meyer avoid maximum effort for some tasks?
In the figures he cites, higher effort raises estimated cost substantially for a relatively small score increase in some cases. His example is Opus 5.5 at max, estimated at $5.98 per task versus $1.82 at high.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
