🤖

NotTooOldForAI Draft Review

Article draft, password required

Wrong password. Try again.

NotTooOldForAI: Draft

I asked 7 AIs to destroy my business idea.

Here's what each one caught, what each one missed, and what you can steal from this for your own decision.

By Annette Thompson  |  June 2026  |  ~2,200 words  |  NotTooOldForAI

📝 Draft for Annette's review. Ready for your edit pass before publishing to Substack.

Three weeks ago, I was genuinely excited about a business idea.

It had everything I wanted: a social mission I believed in, a tech stack I could build myself, and a name I loved. I'd been sketching it out for weeks. I'd done my research. I was close to starting.

What I didn't have was anyone to tell me what was wrong with it.

So I did something that most business advisors would never think to suggest: I ran the entire concept through seven AI models at the same time, gave all of them the same brief, and told every single one of them to try to kill it.

🎬 GIF PLACEHOLDER: someone confidently walking into a room, then realizing it's a courtroom with judges on all sides

The result wasn't what I expected. Different models caught different things. One found a contradiction in the mission itself. Another found a risk hiding in the payment processor fine print. A third questioned a feature I was most proud of.

Together, they gave me a board of advisors I never could have assembled any other way.

New here? Welcome.

I'm Annette. I'm 57, learning AI in public, and this newsletter is for everyone who felt like the AI train left without them. It didn't.

Subscribe free

🤖 Why adversarial review works (and why most people don't use it)

Most people use AI like a collaborative partner. They ask it to help them build something, and the AI helps them build it. ChatGPT, Claude, Gemini: they're all optimized to be helpful. That can make them a little bit sycophantic.

You describe your idea. The AI nods along. It helps you flesh out the details. It never says "have you considered that this might be illegal in your state?"

Adversarial review flips the frame. Instead of "help me build this," you say: "Tell me everything that could go wrong. Be the skeptical investor who would never fund this. Find the fatal flaw."

📊 DIAGRAM NEEDED: simple side-by-side, "Standard AI prompt: help me build this" vs. "Adversarial prompt: try to kill this"

Engineers use this technique to audit code before they ship it. Security teams use it to find vulnerabilities before hackers do.

But here's the insight that changed how I work: it's just as powerful for business ideas, pricing strategy, partnership terms, and any high-stakes decision where you want the uncomfortable truth before you invest real time and money.

And the reason to use MULTIPLE models instead of one? Each AI has different training data, different reasoning patterns, and different blind spots. One model will catch what another one misses. Running seven isn't paranoia. It's assembling a diverse advisory board for the cost of a cup of coffee.

💡 The business idea I put on trial

The concept was called BackOfficeStack: a bilingual software platform to run home-service businesses in Boulder, Colorado, starting with house cleaning. One design choice set it apart. Built right into the code: a living-wage floor. Any booking where a worker's take-home would drop below $26 per hour gets blocked before it can proceed.

The mission: use software to eliminate admin overhead so more of every customer dollar reaches the worker instead of a middleman.

I thought it was clever. I thought the ethics were airtight. I thought the tech was sound.

I wrote it up as a detailed brief: the target market, the unit economics, the tech stack, the legal guardrails already built in, the verticals I'd expand into after cleaning was profitable. Then I handed it to seven models and said: find everything that would kill this.

🔧 The seven-model firing squad

I ran all seven in parallel using a Python script and a service called OpenRouter (one API that routes to any of these models). Total cost: a few dollars. Total time: about fifteen minutes.

Here's what each one brought to the table.

Claude Opus 4 - deep strategic thinking

Claude didn't just find flaws. It found the flaw that made all the other flaws worse.

"The mission and the guardrails contradict each other. You want to maximize jobs for Mexican workers but explicitly refuse to verify work authorization. That's not a clever workaround. That's structuring a business to provide labor to an undocumented workforce while maintaining plausible deniability."

That's not a legal technicality. That's a mission-level contradiction. Claude identified that the philosophical problem was built into the foundation itself, not just a compliance gap.

Claude also did the most nuanced unit economics, factoring in self-employment tax (which cuts take-home by about 15%), drive time between jobs, and supply runs. Its conclusion: to deliver a true $26/hr after all expenses, the average cleaning would need to be priced at $250-$300. That's the top of the Boulder market.

Best at: Systems thinking, philosophical coherence, finding the contradiction that everything downstream flows from.

o3 - structured reasoning, finds logical flaws

OpenAI's reasoning model approached the review like a legal deposition. Clean numbered lists. Clear logical chains. Specific statutory citations.

"Colorado applies the strict ABC test: the worker must be free from control, perform work outside the usual course of business, and be customarily engaged in an independent trade. You fail all three. Penalties: unpaid payroll taxes, 0.5 to 20 percent interest, treble damages for willful withholding."

Where Claude worked in narrative insight, o3 gave the flowchart: premises, consequences, verdict. It also gave the most actionable "what could actually work" section: four numbered alternatives, each concrete and legally cleaner than what I'd proposed.

Best at: Logical precision, specific legal detail, clear and actionable recommendations.

Gemini 2.5 Pro - systematic market analysis

Gemini cut straight to the immigration problem within its opening paragraph, faster than any other model:

"The mission to 'create the maximum number of living-wage jobs for Mexican workers' combined with the policy of 'no work-authorization document collection' creates a direct and unavoidable legal catastrophe. This isn't a clever workaround. It's willful blindness."

Gemini's framing was sharper on this specific point than the others. While Claude took two paragraphs to build the argument, Gemini stated it immediately as a categorical fact. It's particularly good at recognizing whether a concept matches known failure patterns.

Best at: Fast pattern recognition, decisive legal framing, calling a fatal flaw a fatal flaw without hedging.

Grok 4 - adversarial and contrarian by nature

Grok took the most interesting angle on the unit economics. Rather than immediately showing why the numbers fail, it first showed a scenario where they DON'T fail, then revealed why that scenario never actually happens.

"A realistic operations cut that preserves $26 net after taxes requires platform fees below 8%, which leaves the business unable to fund marketing, support, or scaling."

That's more sophisticated than "the math doesn't work." It's: the math CAN work, but only at a platform margin so thin that you can't run the business. The concept starves itself of operating capital in order to honor its own mission.

Best at: Finding the non-obvious version of the problem. The "yes, but here's why that doesn't save you" move.

GPT-4o - solid business acumen

GPT-4o gave what a traditional business consultant would give: organized, professional, readable, covering all the bases. It identified the same core issues as the other models.

Where it was weaker: the unit economics section skipped the self-employment tax calculation, and the "what could actually work" recommendations were more generic. "Leverage her SEO expertise to lower customer acquisition costs" isn't wrong. It's just not surprising.

Best at: Business structure, accessibility, clean framing. Good for a first-pass overview before going deeper with other models.

DeepSeek V4 - strong reasoning, different priors

DeepSeek found a risk that every other model completely missed. Every other model focused on legal risks with workers, customers, and regulators. DeepSeek thought about the payment processor.

"Stripe's terms of service require you to disclose your business model and may flag a platform that facilitates payments to workers whose work authorization status is unknown. Stripe has compliance teams that review accounts. If they determine you are facilitating payments to potentially unauthorized workers, they will freeze your account and hold your funds for 90 to 180 days."

If your payment processor freezes your account, your business is dead regardless of everything else you've gotten right. DeepSeek also cited the actual Colorado cooperative statute (CRS 7-56-101), not just a general warning about complexity.

Best at: Finding the second-order risk one layer below where everyone else is looking. Infrastructure blind spots.

OpenRouter Owl Alpha - high-usage general reasoning

Owl Alpha caught something none of the other six mentioned. I'd been proud of the bilingual feature. English and Spanish, from day one. It felt like part of the mission.

"Bilingual from day one doubles your surface area for bugs, support tickets, and customer confusion without doubling your revenue. Every email template, every error message, every Stripe receipt, every calendar invite must be maintained in two languages. For a solo operator with no home-services experience, this is a massive tax on iteration speed."

It also gave the most grounded advice: "Hire two cleaners as W-2 employees, get insurance, build a simple website, and start knocking on doors. The software can come later, after you understand what the business actually needs."

Best at: Operational thinking. Keeping asking "but who actually does this, and what does it cost them?"

📊 What all seven agreed on

The consensus (independently reached)

  • Worker classification is illegal under Colorado's ABC test. If the platform sets the price, blocks bookings, and controls the customer relationship, the workers are employees. All seven said this.
  • The unit economics break at the low end of the market. At $150-180 per cleaning, there isn't enough margin to pay $26/hr AND run a sustainable business. All seven showed different math, same conclusion.
  • The cooperative structure is 6-12 months of legal work, not a shortcut. Articles of incorporation, bylaws, member agreements, securities compliance. All seven mentioned this.
  • No home-services operations experience is a real gap. Software doesn't replace knowing what to do when a cleaner no-shows at 9am.

🎯 The things ONLY ONE model caught

This is the part that changed how I think about this technique.

If I'd run just one model, I'd have gotten a good review. But I'd have missed:

  • Only DeepSeek caught the Stripe account freeze risk. Your payment processor can kill you before a regulator even opens a file.
  • Only Owl Alpha flagged the bilingual maintenance overhead. A feature you love is still a liability if you can't maintain it at your current team size.
  • Only Claude identified that national-origin targeting is itself a potential civil rights problem, separate from the immigration question.
  • Only Grok showed the 8% platform margin trap, where the math "works" in one scenario and fails in all the adjacent ones.
Running one AI is getting a second opinion. Running seven is running a stress test.

🔧 How to do this yourself

You don't need a Python script or an API account. Here's the free version:

1
Write your idea in 3-4 paragraphs. Include: what it is, who the customer is, what you'll charge, what your costs are, and what relevant experience you have (or don't).
2
Open Claude, ChatGPT, AND Gemini in three separate browser tabs. All three have free tiers. All three have meaningfully different reasoning patterns.
3
Paste this exact prompt into all three:
Here is my business concept. Act as an adversarial business analyst. Find every fatal flaw, questionable assumption, and legal risk. Be direct and specific. Structure your response: 1. Fatal flaws that would kill this outright 2. Unit economics, show the actual math 3. Competitive risks: who's already doing this? 4. Legal and regulatory exposure 5. Questionable assumptions 6. What could actually work
4
Read all three side by side. Look for where they disagree. Look for the thing only one of them caught. The disagreements are as interesting as the agreements.
🎬 GIF PLACEHOLDER: three documents appearing side by side, each highlighting different sections in different colors

The most important thing: don't just ask one model. The value is in the differences.

What I did with the results

I didn't abandon the project.

I came out of the review with a much clearer picture of what had to change first: the cooperative structure needs a real attorney. The unit economics need to be tested at $250+ price points. The bilingual feature belongs in recruiting, not customer-facing flows, until the operations are stable.

That's almost nothing in cost for input that saved me from building in the wrong direction for months.

The right advisor doesn't always tell you what you want to hear. Seven of them almost never will. That's exactly why they're worth consulting.

Send this to the friend who's been polishing a business plan instead of stress-testing it.

If you found this useful:

Subscribe free to get every issue of Not Too Old for AI. Then reply and tell me: what's the idea you've been afraid to put in front of a firing squad?

Subscribe to NotTooOldForAI

Annette Thompson writes Not Too Old for AI for women and men 50+ who are building something new with AI. She does it out loud, with receipts.