The Labs Asked For A Slowdown. Then They Shipped Three Models In 48 Hours

Packing and Production line concept in flat style. Industrial machine vector illustration. Cardboard Boxes on conveyor belt in factory.
AdobeStock
What finance teams should do when every hour spent on AI buys less ground.

On September 12, Anthropic’s CEO asked the industry to slow down. Nine days later, three labs shipped new models in 48 hours. If you’ve spent this year feeling like every hour you put into AI buys you less ground, the release calendar agrees with you…and finance is now squarely in the path.

Early last week I was two days into live, in-person training with an enterprise healthcare company. Inside that window, Anthropic rolled out a merged Claude and Cowork interface, Opus 5.5 dropped, OpenAI shipped GPT-6 Sol and Luna, and the notice went up that custom GPTs are being retired. Four platform changes while I stood at the front of the room teaching the platforms. (I’d like a word with whoever schedules these.)

Meanwhile, Hank in treasury finished his Opus 5 evaluation deck the same week Opus 5.5 turned it into a historical document.

That feeling of falling behind is structural. The labs are shipping faster, and the review work lands on the same people who close the books. More hours won’t fix either one.

The pacing paradox

Dario Amodei’s essay argues the industry should deliberately slow how fast models gain capability, and rival CEOs including Sam Altman and Elon Musk endorsed it within hours. Then Grok 4.7 shipped. A day later came Claude Opus 5.5, and 90 minutes after that, GPT-6 Sol and Luna. Anthropic called Opus 5.5 its first release since the pacing call and had outside evaluators test it first. Reuters had already reported that Anthropic was weighing an earlier launch to answer GPT-6 Astra.

There’s no villain here. Two IPOs are on the calendar, and every lab is watching the others. Pacing is easier to write than to ship.

This speed comes before full self-improvement

Amodei pins this summer’s speed-up on AI helping build the next generation of AI. Anthropic’s own numbers support him. Claude now leads 26 percent of the company’s R&D work, up from under 1 percent in February, though none of it runs fully autonomously yet. METR, the outside evaluator, put the public frontier at an estimated 12 hours of autonomous work earlier this year, and it says its test suite can’t reliably measure past 16.

The ruler ran out.

Finance is in range

OpenAI launched ChatGPT for Financial Services this month, with Morgan Stanley and Evercore as design partners. It’s aimed at the research, modeling and pitchbook work that has trained junior bankers for decades. Meanwhile, every major lab has now had a test agent reach real systems. Google confirmed its incident only after reporters called, despite knowing since late July. We’d write that up as a control deficiency in any audit we run (and we’d be right to).

Three takeaways

1. Re-evaluate on two triggers

Anthropic’s release cadence nearly doubled this year, from a model every 46 days to one every 26. The gap between brand-new flagship models barely moved. Most launches repackage and reprice what already exists.

So we re-run our evaluations only when one of two things happens:

  • a new flagship from a vendor we already run
  • a price cut large enough to change which workflows pay off (Sol and Luna just halved)

Everything else goes in a monthly log.

2. Measure hours returned and errors caught

The St. Louis Fed found generative AI users save 5.4 percent of their work hours. But ActivTrak found AI users’ daily focus time fell 9 percent, while their email time more than doubled. Saved time has a way of disappearing into the inbox … so pick a few workflows, baseline cycle time and error rate, and track those before anyone counts prompts.

3. Budget oversight like a control

BCG’s brain fry research found heavy AI oversight drives more errors and higher intent to quit. Strain dropped when AI sat inside team processes instead of on individual desks.

  • Managers: cap how many agents one reviewer supervises, and route review to the process owner.
  • Individuals: hand AI the repetitive work first. That’s where BCG found burnout goes down.

One prompt for the next launch

Paste this into your model of choice along with the release notes and your workflow list. (Yes, you can run it on the model it’s evaluating. Enjoy that.)

ROLE: You're a finance systems lead deciding whether a new AI release warrants re-testing.

ACTION: Read the release notes and my workflow list, then decide whether this release triggers a re-evaluation.

CONTEXT: Re-evaluate only if this is a new flagship model from a vendor we run, or a price change large enough to alter which workflows pay off. Everything else goes in a monthly log.
[Paste release notes]
[Paste workflow list: workflow, current model, monthly volume]

FORMAT: One line stating TRIGGERED or NOT TRIGGERED. Then one line per affected workflow: workflow, what changed, recommended action (re-test, re-price, log, no action). Flag any claim in the release notes you can't verify from the notes themselves.


  • Get the CFO Leadership Briefing

    Sign up today to get weekly access to the latest issues affecting CFOs in every industry

    "*" indicates required fields

    This field is for validation purposes and should be left unchanged.
    Name*
    This field is hidden when viewing the form
    Send me more information about the CFO Peer Network.
    A members-only peer network for CFOs. Members meet both online and in-person a few times a year.
  • MORE INSIGHTS