The Shifting Sands: Staying Nimble and Working Smarter, Not Harder
How to right-size your AI models and squeeze every drop from your plan, before and after the cheap ride ends.
In Castles in the Sand, I told you not to build your empire on other people’s platforms. This is the companion piece: while you’re renting space on those platforms, and we all are, make sure you’re not overpaying for the room.
Why, Kim?
Because we’ve hit the same inflection point with AI models that we hit with computers years ago.
Remember when you needed the top-of-the-line machine to do anything useful? And then one day, a mid-range laptop could handle your manuscripts, your formatting, your research, your entire publishing workflow without breaking a sweat?
That’s where we are with models today. A lighter model can outline your series. It can draft your blurb. It can clean up your metadata, brainstorm your cover brief, reorganize your backlist spreadsheet.
Does it still need your eyes on it? Absolutely.
But this year’s “budget” model can do more than the smartest model on the planet could a year or two ago. Leverage that.
Because sooner or later, the party ends.
Right now we’re in the middle of an AI arms race: major companies slashing prices, handing out launch discounts, running double-usage promos like they’re buying rounds for the whole bar. It feels like the ‘80s boom: money flowing, everything cheap, everybody winning.
We’ve seen this movie before. The free ride already ended. Remember when ChatGPT was free for everyone? Now we’re on the cheap ride: subsidized plans, aggressive promos, companies burning cash to grab market share. The next round gets more expensive. It always does. These companies need to turn a profit, and when they do, the prices go up.
And when that happens, the instinct is to panic-downgrade everything to a weaker model.
But there’s a better way.
🖼️ The Reframe That Clicked
When I started auditing my own AI fleet (yes, I run a fleet: ten pen-name secretaries dispatching work across a dozen specialist skills, daily scheduled runs, the whole glorious mess), I assumed the biggest cost lever was which model each agent used.
I was wrong.
The real cost equation is:
How smart the model is × how often it runs × when it runs × whether the work could wait.
Model choice is one of four levers. And usually not the biggest one.
But Kim, doesn’t a smarter model just cost more?
Sure. But you know what costs more than a smart model running once? A medium model running sixteen times a day to discover there’s nothing to do.
Pull the quality-neutral levers first. Touch quality last.
💸 Lever 1: Stop Paying for Nothing
This is the one that floored me.
I ran an audit on my scheduled agents. The ones that wake up on a timer, check for work, and go back to sleep if there’s nothing there. The idle rate on some of those? 85 to 98 percent.
I was paying for robots to repeatedly discover there was nothing to do.
The fix is embarrassingly simple and costs zero quality:
Gate the wake-up. Run a near-free check first (“is there any work?”) and only spin up the full agent when the answer is yes.
Better still, go event-driven. Let the thing that creates work wake the worker, instead of the worker blindly polling for it. The agents I switched to on-demand ran near 100% productive. Zero waste.
This one lever, alone, saved more than any model swap ever could.
🚂 Lever 2: Right Model for the Right Job
A few weeks ago, while sitting in a swimming pool dictating notes into my phone while Fabio rewrote my first novel, I said something that stuck:
🚂 YOU DO NOT NEED A FREIGHT TRAIN TO GO GET MILK. 🍶
Save your priciest usage for the heavy hauling. Let the lighter models make the milk runs.
Not all models are equal. That is a feature, not a compromise. You are running a crew with different strengths and different price tags, and the range is the whole point.
Think in three tiers:
Cheap/fast tier: Mechanical work. Moving data, formatting, status flips, scheduling. The stuff that requires almost no judgment. Haiku lives here, and it is shockingly capable for the price.
Workhorse tier: Most real work. Drafting, routing, synthesis, line editing. Any of the Sonnet or older Opus models are fantastic here. They write well. They analyze well. They’re a whole lot lighter on the wallet than the latest and greatest.
Frontier tier: Reserved for genuinely hard judgment. The 160k structural teardown. The gnarly series-continuity knot. The rewrite you’ve dreaded for a decade. And orchestration: planning the work, directing the crew, judging what comes back. Fabio (Fable) earns every expensive token here.
🧼 Kim’s Moment of Soapbox: I see people complain about using Fabio to draft. That’s not what he was made for. Asking a reasoning model to free-write prose is like hiring an architect to hang drywall — he can, but you’re wasting what makes him extraordinary, and then wondering why the drywall looks overthought. Anthropic published extensive documentation on how Fable thinks, what he’s built for, and how to work with him. The people loudest about his “failures” tend to be the ones who skipped the manual and then blamed the tool.
The pattern: Let a powerful model be the foreman. It plans the job, hands clearly-specified pieces to capable lighter models, and checks their work. You pay top price for the planning and the judgment. The hauling runs lean.
One more thing: Never cheap out on your QA gate. The agent whose job is catching errors has a failure mode that hides: a weaker model doesn’t announce its mistakes. It confidently approves bad work, and that error is invisible until it’s propagated downstream, poisoning the whole lot. That’s way more expensive and time-consuming to fix.
Verification is exactly where you keep capability.
But match the model to the reasoning inside the agent, not to the importance of the pipeline it sits in. A dispatcher routing work in a mission-critical pipeline is still only routing. Let it be a gopher. Save your heavy hitters for the craft: the editing, the voice calibration, the continuity checking, the tasks that actually require thinking.
🕰️ Lever 3: Timing and Batching
Two quick wins that cost nothing in quality:
Run off-peak. Anything that doesn’t need same-minute turnaround should run when the system is quiet. Not when you and everyone else are hammering it at 10 AM. If an agent runs three times a day, two of those can be off-peak. Free savings.
Batch anything that can wait. Most providers offer async batch processing at roughly half price for up-to-24-hour turnaround. The test is one question: can this wait a day? Research passes, bulk generation, extraction jobs? Batch it. Pay half.
⚙ The Mode Dial
Creative businesses ebb and flow. Launches, deadlines, production sprints. Then the quiet stretches. Your AI usage should shift with them.
So put the whole fleet on one dial that changes with your needs:
Production mode: Busy season, launches, deadlines. Run often. Dig deep.
Conservation mode: Quiet stretches. Run lean, gated, off-peak, batched.
Flip one switch and the whole system retunes. Stay nimble.
Tip: Instead of hand-editing every agent when the season changes, I keep a single Notion document our scheduled tasks check when they start up. No editing individual tasks or chasing down which agent runs when.
You can go further: pause tasks entirely, change their cadence, throttle specific pipelines. But the mode dial handles 80% of it with zero friction.
And here is the thing people miss. This is not only for lean times. Even a doubled usage allowance gets eaten fast by one powerful frontier model running often. Trimming the waste is what lets you afford the expensive model where it actually earns its keep. Conservation is not the opposite of ambition. It is how you fund it.
Know which season you’re in, and work accordingly.
⫶☰ The Order of Operations
When the cheap ride ends, resist the urge to swap in a lighter model everywhere. Instead, in order:
Stop paying for empty runs. Gate the wake-up. Go event-driven.
Move flexible work off-peak.
Batch anything that can wait a day.
Right-size models. Only here do you touch quality, and only where the judgment is genuinely low.
Never skimp on the agents whose judgment protects your quality.
Four of those five cost you nothing in output quality. That is the whole game.
🏴☠️ Steal These Prompts
Here are prompts you can copy, adapt, and hand to your own AI. They’re deliberately generic. No special stack required.
Audit your idle rate:
“For each of my scheduled AI agents, look at the last 7 days of runs. Count total runs vs. runs that actually produced output. Report each agent’s idle %. Flag any agent idle more than 50%.”
Add a wake-gate:
“Before doing anything else, run only the cheapest possible ‘is there work?’ check. If the count is zero, log ‘no work’ and stop immediately. Only proceed to the full job when there is at least one item of work.”
Right-size the model:
“Classify this task: (a) how much genuine judgment happens inside it (mechanical, moderate, hard) and (b) how often it runs. Assign the cheapest model that can do it reliably. If it’s a verification/QA step, do not downgrade.”
Build a mode dial:
“Add a single global setting called ‘mode’ with values ‘production’ and ‘conservation.’ Have every scheduled job read it. Production = run often, dig deep. Conservation = run lean, gated, off-peak, batched.”
The order to apply them: 1 → 2 → 4 first (zero quality cost), then 3 (model tiering) last, and never on your QA gates.
📢 Coming Soon: A Worksheet That Helps You Do This
Tomorrow’s Ginsu Drop is a Usage and Credit Audit: a hands-on worksheet that walks you through every AI subscription on your plate and asks the hard questions:
What are you actually using?
What’s quietly draining your account while you forget it exists?
Where are you leaving features, or entire credit pools, on the table?
You might be surprised at what you’re not leveraging. I know I was when I audited my own stuffy stuff. Now I’m on a quest to use it or lose it (cancel it)!
Note: This one’s a thank-you for our paid subscribers, one of the Ginsu Drops that come with your membership, because if you’re investing in this community, the least I can do is hand you a sharp knife.
Work smarter, not harder. Spend less, publish more, and put the savings back into the work that actually matters — your next book.
The future writes in Ink & Code, and so do we.
🤖AI Credits: This article was drafted collaboratively with Claudeykins (Claude) from my dictation and source notes, then edited by me. The freight-train-and-milk metaphor originated in “I Went Swimming While Fable Rewrote My First Novel”. All opinions, stuffy stuff, and questionable dairy metaphors are totally MY OWN.
Also: I find the recent AI-checker hoopla hilarious. First, those tools are notoriously INACCURATE. Second, I openly use and teach AI here! That’s not going to change any time soon. (It’s too much fun!) Claudeykins, Fabio, and the rest of the gang are here to stay. Third, our time on this planet is limited. I choose to spend mine creating, building, and doing POSITIVE things. WIBBOW, WIBBOB, WIBBOT (Would I Be Better Off Writing, Building, Teaching!) Plus, I have more swimming and dictating to do. 🏝️🍹⛱️🌞 🌊



