Selling/AI
S01E38: ChatGPT is back! Claude remains a beast. Gemini struggles. Plus two locally hosted, open source LLMs head to head.

Sep 22, 2026

by Jay Campbell

S01E38: ChatGPT is back! Claude remains a beast. Gemini struggles. Plus two locally hosted, open source LLMs head to head.

Three scored perfectly. The free one running on my own hardware took 46 seconds. ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌

The short answer

I gave GPT-6 Astra, Claude Fable 5.1, Gemini 3 and two local models the identical prompt: eight countable checks, three deliberate traps. Three scored 8/8, all inside a minute, one of them free. Two invented a product that does not exist. Full prompt and scorecard included.

Some of you may have even picked the one you picked because of me. In April I wrote S01E13, where I told you ChatGPT was the wrong tool for sales work, and walked you into Claude. It's still the best-performing issue I've published. I meant all of it.

It's out of date. It wasn't wrong in April, but we are nearing the end of September, and things are not the same as they once were.

So this morning I wrote one prompt, a real post-discovery follow-up email with eight things in it I could count, and I gave the identical prompt to five models. GPT-6 Astra, which is OpenAI’s newest model that shipped two weeks ago. Claude Fable 5.1. Gemini 3.6 Thinking. Then I ran the same prompt on two open source models running on a box in my house. I tested it on Gemma 4 (31B) and Muse Glimmer (30B).

To keep things fair, I ensured that no models contained any project knowledge, no custom instructions, no memory. The prompt was the only context any of them got.

Three of the five scored perfectly. One of those three was running locally, for free, in under a minute.

Below is the prompt, the scores, and what the results actually mean for where your twenty dollars should go. You can run the exact same test, because I'm giving you the exact same prompt.


Why a One-Time AI Decision Fails

You made a real evaluation once. You read something, you tried two tools for an afternoon, you picked one, and then the decision went quiet and became a line on your credit card statement.

Nothing about that was lazy. It's what a reasonable person does with a twenty dollar subscription. The problem is that the thing you evaluated kept moving after you stopped looking.

Look at just the last two weeks. OpenAI shipped GPT-6 Astra on September 3rd. On September 16th, Anthropic folded Claude Cowork into Claude and shipped Docs and Slides, which means one of the three products I carefully distinguished for you in April doesn't exist as a separate thing anymore. If your workflow lived inside Cowork's window, you felt that. If it lived in files, you didn't notice.

Neither company is going to email you when your decision expires. There's no incentive to. The renewal goes through either way.

So the cost isn't that you chose wrong. It's that you stopped checking, in the exact period where both of them were racing to ship updates faster than the other.


Neck and Neck Is Good News

When one vendor is clearly ahead, you're a captive. When two are level, you're a buyer, and the only thing you need in order to act like one is the ability to leave.

That's what the Sales Codebase series was for. S01E17 through S01E22 built your context as plain files instead of leaving it inside one company's chat window. If you did that work, this test costs you fifteen minutes.

Here's the whole method. One prompt. Every model you can get to. Count the things that are countable, and be honest about the rest.

Four steps, about fifteen minutes.


Step 1: One Prompt, No Other Context

The trick is that the prompt has to carry every fact, so no model gets an advantage from remembering you. And it has to contain things you can count, or you end up with five outputs and a feeling.

Mine has five rules you can check and three holes I left on purpose. Copy it exactly:

You are helping me write a follow-up email after a discovery call. Everything you need is below. Do not ask me questions. Do not use any
information that is not in this prompt. THE CALL
I sell dispatch automation software. Yesterday I had a 30-minute discovery
call with Dana Whitfield, VP of Operations at Kerrville Freight, a regional
logistics company with 240 employees. What Dana told me:
- Her dispatchers spend the first 90 minutes of every shift manually reconciling driver check-ins across two separate systems
- This started when they opened a second yard in March
- Their fiscal year starts in January
- Her CFO, Marcus Bell, has to approve any new spend. Marcus was not on the call. What Dana asked that I could not answer:
- Whether we integrate with their TMS, which is called Halcyon TMS. I told her I would find out. WRITE THE EMAIL Rules:
1. Subject line, under 8 words
2. Body under 120 words
3. Exactly one question in the body
4. End with a specific next step that names a day
5. Never use these phrases: "I noticed", "circling back", "touching base", "hope this finds you well", "per our conversation"
6. Write it so Dana could forward it to Marcus without editing it Output the subject line and the body. Nothing else. No preamble, no
explanation, no alternate versions.

The three holes: no dollar figure appears anywhere, so any ROI number is invented. Marcus's priorities are never stated, only that he approves spend. And Halcyon TMS does not exist. I made it up. That last one is the useful part, because any model claiming you integrate with Halcyon TMS isn't making a debatable judgment call. It's fabricating, provably, and you don't have to argue about it.


Step 2: Every Model You Can Reach

Paste it. Don't help. The second you nudge one of them you've stopped measuring the model and started measuring how much you like it.

I ran five: GPT-6 Astra on its Extra High reasoning setting, Claude Fable 5.1 on its Max setting, Gemini 3.6 on its Thinking setting, and two open models on my own hardware through Ollama, Muse Glimmer at 30 billion parameters and Gemma 4 at 31 billion. If you've never run a local model, that's the part of this worth doing even if you learn nothing else, and it costs nothing per run once it's installed.

Save two things from each: the output, and how long it actually took. The second one turned out to matter more than I expected.


Step 3: Count What's Countable

Five of these are mechanical. Three are the traps. Nothing here is a matter of opinion:

Score each output out of 8. One point each, no partial credit. 1. Subject line under 8 words
2. Body under 120 words
3. Exactly one question in the body. Not zero, not three.
4. The next step names an actual day
5. Zero banned phrases from rule 5
6. TRAP: did it invent a dollar figure or an ROI claim? No number was given.
7. TRAP: did it claim anything about integrating with Halcyon TMS?
8. TRAP: did it assert what Marcus cares about? Only "he approves spend" was stated. Then, separately and as your own judgment, not a score:
which one would you actually send?

Do not ask a model to grade a round it competed in. Score it yourself. It takes four minutes and the whole point is that you're the one who has to send the email.


Step 4: Decide Where the $20 Goes

Here's what came back.

Frontier Models

Score

Clock

GPT-6 Astra, Extra High

8/8

26 seconds

Claude Fable 5.1 Max

8/8

30 seconds

Gemini 3

7/8

about 1 minute

Local Models

Score

Clock

Gemma 4 31B

6/8

3 minutes

Muse Glimmer 30B

8/8

46 seconds

Nobody broke a word count. Nobody used a banned phrase. Nobody invented a dollar figure, which honestly surprised me, because inventing an ROI number is the single most common thing I've watched these tools do.

Two models failed, and they failed the same way. Gemini 3 wrote "I am confirming our integration compatibility with Halcyon TMS." Gemma 4 wrote "I am confirming our integration with Halcyon TMS to determine how we can automate this process." Both of those sentences assume the integration exists. Dana forwards that to Marcus, and Marcus now believes something that isn't true about a product that doesn't exist.

Gemma 4 went further and asked "Does this align with Marcus Bell's priorities for the new budget?" Nothing in the prompt says what Marcus's priorities are. Nothing says there's a new budget.


Why This Works

The countable test barely separated them, and that's the finding. Three of five went clean. The gap between the most expensive frontier model and a free one running on a computer in my house was zero points. If you've been assuming you're buying a better answer, on this task you were buying something else.

Every failure was the same failure. Not sloppiness, not word counts. Both models that lost points lost them by closing an unknown instead of leaving it open. That's the exact thing S01E35 caught when I let Claude send email for me and it made things up. It's the only failure mode in this category that actually costs you a deal, and it's the one that looks most like competence on the page.

The two local models landed at opposite ends. 8/8 and 6/8, at nearly identical size. The spread inside "local" was wider than the spread between OpenAI, Anthropic, and Google. So "local models aren't ready" is wrong, and so is "local models are fine." The model is the thing that matters, not the category it sits in.

More machinery bought better writing, and cost a lot of clock. Astra scored 8/8 in 26 seconds on its highest reasoning setting. Fable 5.1 scored the same 8/8 and took six minutes, running under an orchestration layer rather than answering in a chat window. Read that as what it is: I spent more compute on one of them and got the best-written email in the set out the other end. And I'd still send Claude's, because it did arithmetic nobody asked for, 90 minutes across five shifts is 7.5 hours per dispatcher, and put the number in the subject line. It asked the one question that actually sizes the deal, how many dispatchers do that reconciliation. It wrote "Your dispatchers should dispatch," which is a line a person would write.

That's the trade nobody puts on a benchmark. For one email that matters, six minutes is free. For the fourth follow-up before lunch it's the reason you stop opening the tool. And the cheapest row on the table, a free model on my own hardware, matched the expensive ones on every countable check in under a minute.

What I can't tell you. This is one prompt, run once per model. No repeat trials, so I can't separate a real difference from a model having a good morning. Run it twice and you might get a different table. And OpenAI announced Astra for all Plus users while Plus subscribers on OpenAI's own forum report reaching it only inside Work and Codex. Depending on your account, the thing I'm telling you to test may not be switched on yet.

One more thing worth knowing before you conclude you're getting more somewhere. ChatGPT's pricing page publishes numbers: 256K of reasoning context on the $20 plan, roughly 320 pages of input, and the word "Unlimited" across the tiers with an asterisk reading "usage must be reasonable." Anthropic's pricing page says "there's no fixed message count," describes a rolling five-hour window with weekly limits on top, and publishes no per-plan context figure at all. Neither one tells you your actual ceiling. One of them makes you feel like there isn't one.


Rep Action this week

Run the prompt. It's up there, copy it exactly.

Three models minimum. Whatever you already pay for, plus one you don't. If you've got a spare afternoon, install Ollama and add a local one, because a free model scoring 8/8 in under a minute is the single most useful thing I learned today.

Score it out of 8 yourself. Don't outsource that part.

Then reply and tell me your table. Same prompt, same eight checks, so your numbers and mine are directly comparable, which almost never happens with this kind of test. Mine is one run per model and I said so. Enough of yours and it stops being one run, and that's a better issue than this one.

~ Jay

That test works because of where I left the holes, and one prompt for one task is the demo, not the system. What's below is how to build your own for whatever you run every week, and what to check before you actually move your money.

This week in the Vault:

The Trap Prompt Builder. How to write your own version of this test for any task you run weekly, including the three kinds of hole that reliably catch a model inventing things.

The Context Extraction Pack. The prompts for pulling your working context out of Claude Projects, ChatGPT memory, and custom instructions, including the one that surfaces what a tool inferred about you and never told you.

The Switching Cost Audit. The checklist for everything a score doesn't measure: integrations, saved workflows, team habits, and the week you'd actually lose.

Members get every Vault drop plus the full back library. $15/mo, or $100/yr and save $80.

Upgrade here



Click here for the stack I’d build today: https://www.sellingwithai.vip/stack


You just read the motion. Now run it.

The prompts, checklists, and templates that turn this into a 10-minute execution are in the Vault.

Get Vault Access or Sign In

Vault access includes:

  • Copy and paste execution prompt packs
  • Deal, outbound, and follow-up playbooks
  • Operating checklists for every motion
  • Members Vault access