Every AI application startup calls someone else's model — OpenAI's, Anthropic's, or an open-weights stack — and the market's lazy shorthand calls all of them wrappers. The documented reality is more discriminating: some API-calling companies built durable businesses, and thin orchestration layers died when the labs shipped their features. The distinction is testable, and this is the test — five questions from the documented record of who survived. Honey Badgers publishes information, not business advice.
Question one: what happens when the model gets better?
The wrapper test's sharpest edge. A thin wrapper gets worse when models improve: the model absorbs the orchestration, the prompt engineering, the output structuring that was the product. An AI-native application gets better — the model's improvement lands on a product whose value compounds with capability. The documented case: coding assistants — every model-generation improvement increased the value of the surrounding product, because the product was the workflow (editor, repository context, review), not the completion. The counter-case: the dozens of 2023 writing-tool wrappers whose summarization and drafting layers became one-line features of the models themselves within eighteen months. If a model release is bad news for your roadmap, you have your answer.
Question two: do you own data the loop needs?
The second question: is there proprietary data in the product's loop — proprietary either because you collected it, you synthesized it from usage, or your customers will not let anyone else have it? Model providers train on the public internet; they do not train on your workflow telemetry, your vertical corpora, or your customer's private context. The documented survivors of feature absorption owned a data asset: vertical GPT companies with domain corpora, enterprise tools whose context window is the customer's own system of record. A useful stress test: could a competitor replicate your product this quarter, with your permission, using public models and your feature list? If yes, the data moat is zero.
Question three: where does distribution come from?
Feature absorption kills undifferentiated products; distribution decides which differentiated ones monetize. The documented patterns of durable AI applications: products embedded where work already happens (the IDE, the CRM, the spreadsheet), products with their own demand brand (consumer tools that became verbs), and products whose buying center is regulated procurement that punishes switching. The wrapper failures shared the opposite: SEO-dependent acquisition that model providers' answer engines are actively absorbing — the documented traffic declines across search-dependent content businesses through 2024-2025 are the same force pointed at SEO-dependent AI tools.
Question four: is your margin structure a feature or a bug?
The economics question the 2024-2025 record forced: inference costs of 20-40 percent of revenue are survivable only with pricing power, and pricing power comes from the first three questions. The wrapper margin structure — API cost passed through at a thin markup — is structurally doomed in both directions: model price cuts attract competitors, and model improvements absorb the product. The AI-native margin structure — inference as a known unit economic, priced into the product's value metric — survived. The test in one number: gross margin after inference, tracked quarterly. If it is falling as volume grows, the supplier owns the economics; if it is flat or rising, the product does.
Question five: what did you ship that the labs won't?
The last question is about intent: the labs ship what serves their frontier race — general capabilities, broad-appeal features, platform-level tools. They demonstrably do not ship deep vertical workflow, compliance surface, liability acceptance, or unglamorous integrations with the systems where specific industries actually run. The documented survivors are products whose core is something the labs are structurally unlikely to build: the audit trail for insurance claims, the HIS integration for hospitals, the regulatory filing surface. 'The labs won't do this' is a real answer when it is backed by the lab's economics — and a cope when it is backed by hope.
What does a passing score look like?
A durable AI application answers at least three of the five with specifics: a capability that compounds with model improvement, a data or integration asset, a distribution position, a margin structure that survives pricing cycles, or a core the labs' economics will not fund. The wrapper insult should be retired for the businesses that pass — and internalized honestly by the ones that do not, because the market's verdict on thin layers has been consistent, documented, and final.
Calling an API was never the sin. Believing the call was the product was — and the five questions catch it early enough to fix.
For more context, read Six Signals a Startup Should Pivot Before the Money Runs Out.
For more context, read startup postmortems lessons.
For more context, read Acquihires: How Talent Deals Get Priced and Who Gets Paid.

