I asked 7 AIs to destroy my business idea.
Here's what each one caught, what each one missed, and what you can steal from this for your own decision.
Three weeks ago, I was genuinely excited about a business idea.
It had everything I wanted: a social mission I believed in, a tech stack I could build myself, and a name I loved. I'd been sketching it out for weeks. I'd done my research. I was close to starting.
What I didn't have was anyone to tell me what was wrong with it.
So I did something that most business advisors would never think to suggest: I ran the entire concept through seven AI models at the same time, gave all of them the same brief, and told every single one of them to try to kill it.
The result wasn't what I expected. Different models caught different things. One found a contradiction in the mission itself. Another found a risk hiding in the payment processor fine print. A third questioned a feature I was most proud of.
Together, they gave me a board of advisors I never could have assembled any other way.
New here? Welcome.
I'm Annette. I'm 57, learning AI in public, and this newsletter is for everyone who felt like the AI train left without them. It didn't.
Subscribe free🤖 Why adversarial review works (and why most people don't use it)
Most people use AI like a collaborative partner. They ask it to help them build something, and the AI helps them build it. ChatGPT, Claude, Gemini: they're all optimized to be helpful. That can make them a little bit sycophantic.
You describe your idea. The AI nods along. It helps you flesh out the details. It never says "have you considered that this might be illegal in your state?"
Adversarial review flips the frame. Instead of "help me build this," you say: "Tell me everything that could go wrong. Be the skeptical investor who would never fund this. Find the fatal flaw."
Engineers use this technique to audit code before they ship it. Security teams use it to find vulnerabilities before hackers do.
But here's the insight that changed how I work: it's just as powerful for business ideas, pricing strategy, partnership terms, and any high-stakes decision where you want the uncomfortable truth before you invest real time and money.
And the reason to use MULTIPLE models instead of one? Each AI has different training data, different reasoning patterns, and different blind spots. One model will catch what another one misses. Running seven isn't paranoia. It's assembling a diverse advisory board for the cost of a cup of coffee.
💡 The business idea I put on trial
The concept was called BackOfficeStack: a bilingual software platform to run home-service businesses in Boulder, Colorado, starting with house cleaning. One design choice set it apart. Built right into the code: a living-wage floor. Any booking where a worker's take-home would drop below $26 per hour gets blocked before it can proceed.
The mission: use software to eliminate admin overhead so more of every customer dollar reaches the worker instead of a middleman.
I thought it was clever. I thought the ethics were airtight. I thought the tech was sound.
I wrote it up as a detailed brief: the target market, the unit economics, the tech stack, the legal guardrails already built in, the verticals I'd expand into after cleaning was profitable. Then I handed it to seven models and said: find everything that would kill this.
🔧 The seven-model firing squad
I ran all seven in parallel using a Python script and a service called OpenRouter (one API that routes to any of these models). Total cost: a few dollars. Total time: about fifteen minutes.
Here's what each one brought to the table.
Claude Opus 4 - deep strategic thinking
Claude didn't just find flaws. It found the flaw that made all the other flaws worse.
"The mission and the guardrails contradict each other. You want to maximize jobs for Mexican workers but explicitly refuse to verify work authorization. That's not a clever workaround. That's structuring a business to provide labor to an undocumented workforce while maintaining plausible deniability."
That's not a legal technicality. That's a mission-level contradiction. Claude identified that the philosophical problem was built into the foundation itself, not just a compliance gap.
Claude also did the most nuanced unit economics, factoring in self-employment tax (which cuts take-home by about 15%), drive time between jobs, and supply runs. Its conclusion: to deliver a true $26/hr after all expenses, the average cleaning would need to be priced at $250-$300. That's the top of the Boulder market.
o3 - structured reasoning, finds logical flaws
OpenAI's reasoning model approached the review like a legal deposition. Clean numbered lists. Clear logical chains. Specific statutory citations.
"Colorado applies the strict ABC test: the worker must be free from control, perform work outside the usual course of business, and be customarily engaged in an independent trade. You fail all three. Penalties: unpaid payroll taxes, 0.5 to 20 percent interest, treble damages for willful withholding."
Where Claude worked in narrative insight, o3 gave the flowchart: premises, consequences, verdict. It also gave the most actionable "what could actually work" section: four numbered alternatives, each concrete and legally cleaner than what I'd proposed.
Gemini 2.5 Pro - systematic market analysis
Gemini cut straight to the immigration problem within its opening paragraph, faster than any other model:
"The mission to 'create the maximum number of living-wage jobs for Mexican workers' combined with the policy of 'no work-authorization document collection' creates a direct and unavoidable legal catastrophe. This isn't a clever workaround. It's willful blindness."
Gemini's framing was sharper on this specific point than the others. While Claude took two paragraphs to build the argument, Gemini stated it immediately as a categorical fact. It's particularly good at recognizing whether a concept matches known failure patterns.
Grok 4 - adversarial and contrarian by nature
Grok took the most interesting angle on the unit economics. Rather than immediately showing why the numbers fail, it first showed a scenario where they DON'T fail, then revealed why that scenario never actually happens.
"A realistic operations cut that preserves $26 net after taxes requires platform fees below 8%, which leaves the business unable to fund marketing, support, or scaling."
That's more sophisticated than "the math doesn't work." It's: the math CAN work, but only at a platform margin so thin that you can't run the business. The concept starves itself of operating capital in order to honor its own mission.
GPT-4o - solid business acumen
GPT-4o gave what a traditional business consultant would give: organized, professional, readable, covering all the bases. It identified the same core issues as the other models.
Where it was weaker: the unit economics section skipped the self-employment tax calculation, and the "what could actually work" recommendations were more generic. "Leverage her SEO expertise to lower customer acquisition costs" isn't wrong. It's just not surprising.
DeepSeek V4 - strong reasoning, different priors
DeepSeek found a risk that every other model completely missed. Every other model focused on legal risks with workers, customers, and regulators. DeepSeek thought about the payment processor.
"Stripe's terms of service require you to disclose your business model and may flag a platform that facilitates payments to workers whose work authorization status is unknown. Stripe has compliance teams that review accounts. If they determine you are facilitating payments to potentially unauthorized workers, they will freeze your account and hold your funds for 90 to 180 days."
If your payment processor freezes your account, your business is dead regardless of everything else you've gotten right. DeepSeek also cited the actual Colorado cooperative statute (CRS 7-56-101), not just a general warning about complexity.
OpenRouter Owl Alpha - high-usage general reasoning
Owl Alpha caught something none of the other six mentioned. I'd been proud of the bilingual feature. English and Spanish, from day one. It felt like part of the mission.
"Bilingual from day one doubles your surface area for bugs, support tickets, and customer confusion without doubling your revenue. Every email template, every error message, every Stripe receipt, every calendar invite must be maintained in two languages. For a solo operator with no home-services experience, this is a massive tax on iteration speed."
It also gave the most grounded advice: "Hire two cleaners as W-2 employees, get insurance, build a simple website, and start knocking on doors. The software can come later, after you understand what the business actually needs."
📊 What all seven agreed on
The consensus (independently reached)
- Worker classification is illegal under Colorado's ABC test. If the platform sets the price, blocks bookings, and controls the customer relationship, the workers are employees. All seven said this.
- The unit economics break at the low end of the market. At $150-180 per cleaning, there isn't enough margin to pay $26/hr AND run a sustainable business. All seven showed different math, same conclusion.
- The cooperative structure is 6-12 months of legal work, not a shortcut. Articles of incorporation, bylaws, member agreements, securities compliance. All seven mentioned this.
- No home-services operations experience is a real gap. Software doesn't replace knowing what to do when a cleaner no-shows at 9am.
🎯 The things ONLY ONE model caught
This is the part that changed how I think about this technique.
If I'd run just one model, I'd have gotten a good review. But I'd have missed:
- Only DeepSeek caught the Stripe account freeze risk. Your payment processor can kill you before a regulator even opens a file.
- Only Owl Alpha flagged the bilingual maintenance overhead. A feature you love is still a liability if you can't maintain it at your current team size.
- Only Claude identified that national-origin targeting is itself a potential civil rights problem, separate from the immigration question.
- Only Grok showed the 8% platform margin trap, where the math "works" in one scenario and fails in all the adjacent ones.
🔧 How to do this yourself
You don't need a Python script or an API account. Here's the free version:
The most important thing: don't just ask one model. The value is in the differences.
What I did with the results
I didn't abandon the project.
I came out of the review with a much clearer picture of what had to change first: the cooperative structure needs a real attorney. The unit economics need to be tested at $250+ price points. The bilingual feature belongs in recruiting, not customer-facing flows, until the operations are stable.
That's almost nothing in cost for input that saved me from building in the wrong direction for months.
Send this to the friend who's been polishing a business plan instead of stress-testing it.
If you found this useful:
Subscribe free to get every issue of Not Too Old for AI. Then reply and tell me: what's the idea you've been afraid to put in front of a firing squad?
Subscribe to NotTooOldForAIAnnette Thompson writes Not Too Old for AI for women and men 50+ who are building something new with AI. She does it out loud, with receipts.