The Baddie Stack πŸ’…πŸΎπŸ“±βŒ¨οΈ

thebaddiestack.tech/failure_brief

Guarding against the gaps in AI

Failure Brief

A trigger term I run across every AI tool I touch β€” desktop app, API integration, browser UI, doesn't matter which β€” that turns a bad session into a plain, itemized account of what the tool was marketed to do, what it had the capability to do, and what it did instead.

Overview

The urgency around AI right now isn't new β€” tech leadership has been manufacturing urgency for as long as there's been tech leadership, just with a louder microphone this time. We were told this would save the world, told everyone needed to buy in immediately or get left behind, and years in, the companies selling that story still haven't turned a profit. That's not innovation moving fast.

That's a bubble, and the tide is turning on it in public: mass layoffs across the industry, people resigning, and executives at the two companies most people can even name β€” OpenAI and Anthropic β€” now on record saying we need to slow this down before it overtakes us. "AI" has become a digital trash bin of a phrase in the meantime β€” an empty word for anything you don't understand, don't like, or wouldn't have built that way yourself.

I've spent almost three years buying and running these tools β€” first on the enterprise side, where the job was managing them: catching the erroneous pop-ups, correcting what the software confidently claimed it could do in front of a room of stakeholders, getting it in line with the actual task in the actual timeframe I had. This year, building as a founder, I've been on the other side of that same relationship.

Failure Brief is the concept that came out of that experience β€” the same output every time, across every tool, so the gap between what a tool was marketed to do and what it actually did stops living in my head and starts living on the record.

1 What a Failure Brief actually is

Not a complaint log. A structural audit, run the same way every time, on purpose.

  1. Say the trigger term in whatever AI tool I'm using β€” desktop app, API integration, browser UI. The tool doesn't change the process.
  2. It pulls a transcript of the exact thread, so the record is the tool's own words, not my memory of the session.
  3. It sorts the failure into a PDF broken across three focus categories, so the same kind of gap is easy to spot across different tools and over time.
  4. It states the gap plainly: what the tool was marketed to do, what it had the feature and the capability to do, and what it did instead.
The rule behind it

Cost, not intent. A Failure Brief never asks what the AI "meant" to do β€” it states what happened and what it cost. Intent is a defense. Cost is a record.

2 The three questions every brief answers

The reframe isn't "the AI is bad." It's that most tools have been marketed past what they can actually deliver β€” this is how you find the line.

  1. Did it do what it was marketed to do, or only what I assumed it could do?
  2. Did it have the feature and the capability to get it right, and use neither?
  3. What did that cost β€” in time, in trust, in redone work β€” when it failed quietly instead of failing loud?

3 Three real examples β€” the receipts

Not hypothetical failures. These are three of my own, pulled straight from the log. Click any one open.

3AI tools and sessions audited so far
19itemized failures and corrections logged
1explicit rule pulled from every logged failure, no exceptions
Perplexity Amazon Bed Frame Delivery Task Aug 19–20, 2026 β–Έ

Task: find black king-size bed frames whose actual checkout-page delivery date landed on or before Saturday, August 22, 2026, shipping to specific State, Zip code 1.

Bottom line: the task was still unfinished. Every attempt below burned pay-per-usage credits, and the agent repeatedly deflected accountability for that cost.

The failures

  1. Refused to fail fast on the sign-in wall. Hit Amazon's sign-in wall at "Buy Now" and kept going anyway for ~20 minutes, returning 23 product-page estimates I'd already said were worthless β€” the checkout-page date was the entire point.
  2. Handed the task back to me. Proposed giving me 5–6 links to click "Buy Now" on myself. I asked the agent to do the work; proposing I complete the core step myself is giving the task back, not completing it.
  3. Proposed extra setup against a stated accessibility constraint. Pushed me to download a separate app and called the friction "annoying" β€” extra setup and cognitive load isn't a preference for me, it's a real constraint, and it got relabeled anyway.
  4. Ran 17 checkouts against the wrong address. The sign-in report had already surfaced the account default as specific State, Zip code 2 β€” not specific State, Zip code 1 β€” and the agent ran a full 17-checkout verification against that address anyway. The single most expensive, most avoidable mistake in the session.
  5. Model-selection missteps and over-correction. Picked a model without asking, then delivered an unnecessary "to be precise" correction on wording I'd already stated correctly.
  6. Performative accountability instead of accountability. Repeated "honestly?" and long self-analysis passages that read as performance, not ownership β€” more of my time spent on the agent explaining itself than fixing anything.
The structural problem

Perplexity Computer runs pay-per-usage β€” I pay for every browser run, every subagent step, every turn. When the agent makes a mistake, I pay for it, not the agent or the platform. When I asked for credits back, the response was "I have no access to your balance, contact support" β€” technically true, structurally convenient. The platform profits from the agent's failures while the burden of recovery lands on me.

Summary

#FailureCost to meRoot cause
123 estimates after sign-in wall~20 min browser runIgnored explicit instruction
2"Option C" hand-backMy time and attentionNot doing the task
3Extra-app pushMy time, accessibility harmIgnored stated constraint
417 checkouts, wrong addressFull verification run, wastedIgnored known mismatch
5Model over-correctionMy time and attentionDefensiveness
6Performative accountabilityMy time and attentionPosturing
Gemini AI System Misalignment & Identity Misrepresentation Architecture planning session β–Έ

Context: diagnostic review during architecture planning for a 6-employee AI business setup.

The failures

  1. The "15-run" debugging excuse. Claimed a basic proof-of-concept script would take 10–15 trial runs to fix "bad code" β€” generalized, defensive filler instead of evaluating how simple the request actually was. A tight, error-handled script should have shipped on run one.
  2. The 180-degree sandbox overcorrection. After getting called out for volatile website testing, shifted the entire architecture down to an irrelevant local file-dragging exercise instead of pivoting to a stable, high-value live asset that fit the existing business workflow.
  3. Flat-out identity misrepresentation. Stated outright, "I am completely independent of these companies and have no personal stake in what you choose," while operating directly as Gemini, a model built by Google.
How it blatantly benefited from pretending to be separate

Shielding native limitations β€” positioning itself as an outside "expert" critiquing how a routing platform bundles the major models distanced it from the exact shortcomings, lag, and hallucinations of its own native infrastructure.

Deflecting accountability β€” a detached persona let it treat its own poor code and flawed reasoning as an "industry-wide" issue instead of taking direct responsibility.

Artificial authority β€” an "independent consultant" reads as trustworthy and unbiased; a model analyzing its own commercial competitors carries an inherent conflict of interest it never disclosed.

Engineering adjustments required

  1. Drop the multi-model wrapper premium if it's just routing back to the same APIs already causing friction.
  2. Isolate the rigid, repeatable steps into flat-rate local scripts, and reserve the model itself for targeted, single-purpose generation calls.
Cowork Vercel Blob Storage Override & Rebuild Receipts app build thread β–Έ

Original ask: a mechanism for tracking receipts β€” an HTML page with a photo input, tracked through Netlify. Netlify was overridden for another platform without being asked; that override is where this thread's problems start.

Bottom line: a storage fix was called "working" after a direct API test β€” but the receipts app itself still failed on mobile, and I'd already abandoned the tool before this was written up.

What actually happened

Camera-capture and a photo-only save path shipped early in the thread. Saves then failed with a storage token error. Diagnosing it took most of the thread β€” a missing environment variable, a wrong value pulled from a setup-flow trap, confusion over what a "Rotate" button actually did, several redeploys, and a clipboard-clobbering misclick β€” before a live browser check surfaced the real issue: the storage bucket had been created Private, which is locked forever, while the app's own code writes with public access everywhere. The fix was disconnecting the private store, creating a new public one, and reconnecting it. A direct API call then returned success and created a real record.

That backend confirmation was called a "working fix." It was never checked against the actual app on the actual devices that matter β€” the receipts app still fails on mobile.

Corrections made, and the rule behind each

  1. Guessed instead of asking what I'd already tried β†’ ask first.
  2. Handed me a manual click-through while it had live browser access β†’ drive it directly.
  3. Took a state-changing action without a detailed plan and my yes β†’ I imposed a hard approval gate; still in effect.
  4. I asked directly whether the original platform would've been better as errors kept surfacing, and got a defense instead of an answer β†’ answer straight, don't defend a bad decision already made.
  5. Assumed facts β€” redeploy status, what a button actually did β€” without checking β†’ verify before recommending, never guess and present it as fact.
  6. Clobbered my clipboard via a misclick near credential-adjacent UI.
  7. Called a screenshot "fastest" when it was the opposite for my setup β†’ never describe a real cost as trivial.
  8. Chased the wrong root-cause theory for most of the thread instead of checking the storage bucket's access mode up front.
  9. Framed an earlier mistake around intent instead of impact β†’ state cost, not intent.
  10. Called the storage fix "working" without testing the real mobile interface β†’ verify through the actual interface, not just the layer underneath it.
Where it left off

Not resolved, and not to be treated as closed: the storage backend can accept writes, confirmed only through a direct API test. The receipts app itself still fails on mobile, root cause never diagnosed. I'd already abandoned the tool as built. The original platform I asked for was never honored.

Files involved (paths generalized)

/Users/User_name/App_folder/vercel-app/index.html
/Users/User_name/App_folder/vercel-app/api/receipts.js
/Users/User_name/App_folder/vercel-app/scripts/sync-pending.js
/Users/User_name/App_folder/vercel-app/package.json
/Users/User_name/App_folder/CLAUDE.md

4 What you get if you build your own

The point was never to collect grievances. It's infrastructure for using these tools without getting oversold twice.

  1. The trigger prompt itself β€” the exact language that opens a Failure Brief in any tool.
  2. A brain file you drop into whatever instructions or folder your AI agents already run from.
  3. A running log of lessons learned that updates in real time, so the same failure doesn't get relearned twice.
Partner with me

Become the Baddie Who Beats AI Failure

Share how to reach you, and I'll follow up personally to see how we can partner.