A FEW THINGS BEFORE WE GET INTO IT
Two Founder Advisory spots are still open for September.
This is the program I run for founders at six figures or early seven figures who are stepping into the actual CEO role and want a real human — not a course, not a community — in their corner for the decisions that come with that. One-to-one. Ongoing. Designed around what you're dealing with, not a fixed curriculum.
If that's where you are, reach out. I'm having short conversations with interested founders over the next few weeks.
WHAT WE’RE DIVING INTO TODAY
Here's a test. Pick one AI workflow your agency runs regularly — writing client content, summarizing research, generating report narratives, drafting pitch angles, whatever you use most.
Now answer this: if the model behind that workflow changed tomorrow — different provider, different pricing, different quirks — could you adapt in a morning?
Not switch permanently. Not rebuild from scratch. Just adapt: update the prompt, run it on the new model, verify the output, and keep working.
If the answer is no, you don't have an AI strategy. You have an AI purchase.
The distinction matters more than most founders realize — because the AI model landscape is not stable, hasn't been stable for three years, and isn't going to be stable anytime soon. The leading model by performance benchmark has changed at least six times in the last eighteen months. Pricing has been restructured, context windows have expanded, models have been deprecated, and new entrants have arrived from directions nobody predicted.
The agencies running an AI strategy handle this as a minor adjustment. The agencies running AI purchases scramble every time something changes — or worse, stay locked to an inferior model because switching feels like too much work.
🗓️ THIS WEEK: The AI Purchase Problem
An AI purchase is what most agencies have. It looks like this:
→ You're using Claude, ChatGPT, or Gemini for specific tasks
→ The people who use it have developed personal prompt habits — not written down anywhere
→ The results are inconsistent across team members because everyone is doing it slightly differently
→ If the model changed, you'd need to re-learn the quirks and rebuild the prompts from memory
→ Nobody has tested whether a different model might do the job better or cheaper
An AI strategy looks like this:
→ Specific use cases are documented as workflows with defined inputs and outputs
→ Prompts are stored in a shared system — not in browser history, not in one person's head
→ The prompt is written to be model-agnostic: it describes what you want, in what format, with what context — without depending on any model's specific personality or training quirks
→ Someone reviews model performance against your actual use cases quarterly — not just checks what's trending
→ Swapping a model on a workflow is a 20-minute test, not a two-week project
The functional test:
Take your three highest-frequency AI workflows. Can you run all three on a model you've never used before, with no ramp-up time, in a single morning? Not perfectly — just functionally?
If yes: you have documented, portable workflows. You have a strategy.
If no: your workflows live in people's heads and model-specific memory. You have purchases.
Why this matters right now:
The model that beats everything else on your specific tasks today is almost certainly not the same model that will beat everything else twelve months from now. That's not speculation — it's the empirical record of the last three years. GPT-4 was the clear leader, then Claude 3 Opus, then GPT-4o, then Claude 3.5 Sonnet, then Gemini 2.0 Flash on certain tasks, then Sonnet 4 and Opus 4. The leaderboard has not stopped moving.
Agencies with a strategy benefit from every model improvement. Agencies with purchases are stranded on whatever they first learned to use.
THE SYSTEM: The AI Strategy Stack
Three components. Every agency running a real AI strategy has all three. Most agencies have none of them, which is how they ended up with purchases instead.
COMPONENT 1 — The Prompt Library
Every prompt your team uses regularly — for any AI task — lives in one shared, searchable place. Not in browser history. Not in a Slack message from six months ago. Not in the head of the one person who figured out how to make it work.
What it contains per entry: the task it performs, the input it expects, the output format it produces, the context or background it needs to work properly, and any model-specific notes (if certain models handle it better). That last field is what makes it portable — the core prompt works anywhere; the model notes are an optional layer on top.
Where it lives: a Claude Project, a Notion database, a structured Google Doc — whatever your team will actually use. The format matters less than the consistency. One place, always updated, accessible to everyone.
COMPONENT 2 — The Workflow Map
For every AI use case in your agency, map exactly where AI sits in the broader workflow: what triggers it, what goes in, what comes out, who reviews it, and what happens next. Not just "we use Claude for content" — specifically which step in which process, with what input, producing what output that feeds what next step.
This is what makes the workflow portable. When someone documents the full step — not just "use AI here" — the specific model becomes a variable you can change, not a fixed dependency. The person following the SOP doesn't need to know how to engineer the prompt from scratch. They run the documented version on whatever model you've designated for that task.
The shortcut that makes this fast:
You don't have to document all AI workflows at once. Start with the three your team uses most often. Those are the ones where inconsistency is currently costing you, and the ones where a model change would hurt most. Document those three this week. The rest follows.
COMPONENT 3 — The Model Review Cadence
Quarterly. Fifteen minutes. Not a full evaluation — a quick benchmark of your top five AI tasks against the current leading models to see if the model you're using is still the right choice.
The process: take the documented prompt from your Prompt Library, run it on two or three current top models, compare the outputs against your quality standard, note the cost difference. That's it. If your current model still wins on your specific tasks, nothing changes. If a different model is now meaningfully better or cheaper, you update the workflow map and move on.
What this does: it turns model evolution from a threat into an advantage. Every time the leading model improves, you benefit from it — because you have the system to capture that improvement without rebuilding anything.
| AI purchase | AI strategy |
|---|---|---|
Prompts live... | In one person's head or browser history | In a shared, versioned Prompt Library |
Model change = | Rebuild everything | Update one field, run the test |
New model releases | "Let's see if it works" (no benchmark) | Tested against your actual use cases |
Team consistency | Everyone does it slightly differently | Documented, auditable, trainable |
Model evolution | A threat — forces reactive scramble | An advantage — you capture improvements |
Institutional knowledge | Leaves when the person does | Stored, documented, transferable |
The AI strategy decisions — what to document first, how to structure the prompt library, how to run the model review, which workflows to build vs. subscribe to — are exactly the kind of high-leverage strategic work that gets done well with a clear-eyed second perspective.
Two Founder Advisory spots open in September. For founders at six figures or early seven figures stepping into the CEO role who want a real human in the room for decisions like these.
→ Express interest in Founder Advisory: [HERE]
⚡THE ACTION: The Morning Swap Test
Run this today. It takes under an hour and tells you exactly where you stand.
Pick your highest-frequency AI workflow.
The one your team (or you) uses most often. Content drafting, research synthesis, report narratives, pitch angles — whatever it is.
Try to run it on a model you don't normally use.
Not switching permanently. Just running the same task on a different model — today, this morning. If you normally use Claude, try GPT-4o or Gemini. If you normally use ChatGPT, try Claude.
Note what breaks and what doesn't.
Does the prompt need significant rewriting? Does the output format shift? Do you need to re-explain context that the other model doesn't have? How long does the adaptation take?
That's your audit result.
If it takes 20 minutes and works well: your workflow is documented enough to be portable. You have a strategy.
If it takes hours or you give up: your workflow lives in model-specific memory. You have a purchase.
The point isn't to switch. The point is to know whether you could.
The ability to swap is the proof of portability. And portability is the proof of strategy.
🤖 AI CORNER: Build a Model-Agnostic Prompt Template
Most prompts fail the portability test not because the idea is model-specific but because the context is implicit. The model you've been using has learned your habits from repeated use — it fills in the blanks from conversational history. A new model has none of that. The prompt that works in one place fails somewhere else because it was never actually complete.
Here's a template structure that makes any prompt model-agnostic — and a Claude prompt to help you rebuild your existing prompts in this format:
PRMOPT:
I want to convert an existing prompt I use regularly into a model-agnostic, documented format that any team member can run on any capable model without losing quality.
Here is my current prompt:
[paste your prompt]
Here is what I typically paste in as context (if anything):
[paste what you normally provide]
Please rewrite it as a complete, self-contained prompt that:
States the role or persona explicitly — don't assume the model knows what kind of agency we are
Defines the input format clearly — what exactly should the user paste in, and in what structure
Specifies the output format exactly — structure, length, tone, what to include and what to omit
Includes any context a brand-new model would need that currently lives in my head or in chat history
Has no model-specific references — it should run on Claude, GPT-4o, Gemini, or any capable model equally
Then give me a one-line title for this prompt that I can use in my Prompt Library, and a 2-sentence description of what it does and when to use it."
Run this on your three most-used prompts. You now have the first three entries in your Prompt Library, each documented well enough that any team member can run them on any model — which is what a strategy looks like.
🛠️ TOOLS OF THE WEEK
Your voice. Every platform. No writing required.
You ghost your own socials by Wednesday. SureThing learns your voice and ships native posts to LinkedIn, X, Instagram, and TikTok, without you writing a thing.
This week's picks are all oriented around model-agnostic AI infrastructure — tools that help you build a strategy rather than a dependency.
OpenRouter openrouter.ai
A single API that routes to 100+ models from Anthropic, OpenAI, Google, Meta, Mistral, and others — with one integration. You write the call once; you change the model by changing one parameter. For agencies building documented workflows: OpenRouter is what makes the model a variable instead of a fixed dependency. Also useful for cost optimization — you can run the same prompt across models and compare price-to-quality. Free to start; pay per token at market rates.
Wordware wordware.ai
An AI workflow builder that treats prompts like code — versioned, testable, deployable, and model-agnostic. Think of it as Notion for AI applications: you build the workflow once, you can run it on any connected model, and you can share it with your team as a tool rather than a raw prompt. Particularly useful for agencies that want to turn prompt-based workflows into proper internal tools without a developer. Paid plans from ~$29/month.
PromptLayer promptlayer.com
Prompt management and observability platform — tracks every prompt run, versions your prompts, logs the inputs and outputs, and lets you compare performance across models on the same task. The model-review cadence from this issue becomes 15 minutes instead of an afternoon when you have PromptLayer running: your historical data shows you exactly how your prompts have performed over time and across models. Developer-friendly; free tier available; paid from $99/month.
Typing Mind typingmind.com
A ChatGPT-style interface that connects to multiple AI providers — Claude, GPT-4o, Gemini, Mistral, and more — in one place, using your own API keys. For founders or team members who want to test prompts across models without switching platforms or accounts, Typing Mind is the fastest option. Also useful for building a shared team AI workspace where different workflows route to different models based on what works best. One-time purchase from $39.
Flowise flowiseai.com
Open-source, visual AI workflow builder for building LLM applications — drag-and-drop interface, model-agnostic, self-hostable. The closest thing to a visual Workflow Map for your AI operations: you define the flow, connect the steps, and can switch out the underlying model without touching the logic. Steeper learning curve than the others but more powerful for complex multi-step workflows. Completely free to self-host.
📊 BY THE NUMBERS
The number that should bother you
If you've been using the same AI model for more than six months without a structured evaluation of whether it's still the right tool for your specific tasks — you've almost certainly been leaving quality or cost savings on the table. Not because your model is bad. Because better options for your specific use cases have probably emerged, and you haven't had the infrastructure to notice.
6 times.
That's approximately how many times the top-ranked general-purpose AI model by benchmark performance has changed since early 2023. GPT-4, Claude 2, GPT-4 Turbo, Claude 3 Opus, GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Flash on specific tasks, Claude Sonnet 4 and Opus 4 — the leaderboard has not held still for more than a few months at a stretch.
The agencies that built their AI operations around documented, portable workflows captured every one of those improvements. They ran their benchmark prompts on the new model, compared outputs, updated the workflow map, and moved on — often in an afternoon. The agencies that built around a specific model's interface and conversational habits treated each change as a disruption.
🔗In Case You Missed It…
AUDIT: Agency AI Value Audit: See if your agency is really AI-native
GUIDE w. prompts: Higgsfield MCP + Flutterflow MCP guides (including what the heck is an MCP) 7 minute read
GUIDE w. prompts: Client Onboarding: The Automation Workflow 9 minute read
GUIDE w. prompts: Claude Managed Agents Build Spec 13 minute read
Take the Operational Debt Scorecard Quiz and see where you’re leaving money on the table
TEMPLATE: Delegation Systems Pack
TOOLKIT: Price Ceiling Toolkit: Part 1 (Pricing Audit Worksheet) and Part 2 (Price Increase SOP with conversation scripts for all three client tiers)
📣 BEFORE YOU GO
Two things:
One thing to do this week: run the Morning Swap Test. Pick your most-used AI workflow, try it on a different model, and time how long the adaptation takes. That number is your current strategy score.
And if you want to work through the AI strategy build with someone in your corner — Founder Advisory spots are still open for September. Reach out to express interest: HERE or reply to this email.
See you next week.
Work smart. Enjoy life harder.
Erin James Murphy
Founder, Agency Owner Lab
When you're ready, here's how we can work together:
→ Agency AI Adoption Assessment + Engagement — Custom AI strategy for your agency. Plus option to add 3 months of fractional ops support to make sure adoption sticks. [Apply here]
→ Agency Growth Roadmap — Operational audit + systems strategy. [Apply here]
→ Founder Advisory — Your advisor. Your business partner. For the founder who uses AI for strategy but wants a real human to strategize with. Quarterly commitments. [Apply Here]
→ Implementation Sprints — done-for-you systems builds, Standalone or paired with another program. [Book a Systems Audit]
→ Agency OS Lab (Community Membership) — SOPs, Claude installs, tool stacks. $97/month. [Join the waitlist here]



