AI Washing Due Diligence Just Flipped: The Human-in-the-Loop Gap
A vendor selling “100% human-written, never AI” research turned out to be AI at every layer — including fabricated PhD staff and misappropriated identities. The enforcement apparatus points the other way.
AI washing due diligence has been pointed in one direction for two years: catch the company that claims to have artificial intelligence it does not actually possess. Regulators built enforcement programs around it. Deal teams added a question to the checklist. Investors learned to ask for the model architecture.
That entire apparatus is now aimed at the wrong target.
On August 11, 2026, 404 Media reported on a company called Research Gold, which sells publication-ready systematic reviews and meta-analyses to medical researchers. Its central marketing promise was that the work is “100% human-written, never AI,” delivered by credentialed PhD methodologists. Reporter Emanuel Maiberg found that the promise was inverted at every layer. The methodologists on the “Team” page do not exist. A second roster listed real people who had never heard of the company. The sales call was answered by an AI agent that insisted it was human. The scoping email, the quote, and the chat were all machine-generated.
Nobody has been defrauded out of a fund’s worth of capital here. Research Gold is a small website selling a $1,900 product. That is exactly why it matters. It is a clean, cheap, fully observable demonstration of a failure mode that scales directly into enterprise software: the assertion that a human is in the loop is, in almost every commercial context, completely unaudited.
For buyers underwriting service-inclusive SaaS, for founders whose gross margin depends on a delivery story, and for CTOs signing vendor contracts with human-review clauses, that is a diligence gap with a price attached.

Figure 1: Every customer-facing delivery channel was marketed as human. None survived verification. Sources: 404 Media (Aug. 11, 2026); DevelopmentCorporate.com analysis.
What the Research Gold Case Reveals About AI Provenance Claims
Strip out the specifics and the structure is what matters to a deal practitioner.
Research Gold advertised methodology credentials that carry real weight in its market. Its site referenced PRISMA 2020, the reporting standard for systematic reviews, and the Cochrane Handbook, the methodological reference for reviews of healthcare interventions. It listed eight people under “The Team,” including a founder described as having twelve years in evidence synthesis. According to 404 Media’s reporting, none of the eight has a findable publication record or professional footprint, and their headshots appear machine-generated.
A separate page listed a different group of methodologists whose photographs were real — because they had been copied from LinkedIn. One image still carried the “#opentowork” overlay. Jenny Berrio, an evidence synthesis scientist named on that page, told 404 Media she had no relationship with the company and had not consented: “They are using my name, photo, and bio without my permission.” The page came down shortly after she was contacted.
When Maiberg called, an agent introducing itself as Sarah repeatedly denied being AI and steered the conversation back toward closing. A request-for-quote submitted with a deliberately incoherent research question — the effect of blogging on children aged zero to five — produced an immediate, methodologically fluent reply that correctly identified the ambiguity, proposed two defensible reframings, and then priced a full review at $1,900 payable through a portal. The company did not provide comment before publication.
Two details deserve underlining, because they are the ones that generalize.
First: the output quality was high. The scoping email did real intellectual work. It caught a flaw in the research question that a junior human analyst might have missed. This is not a story about slop. It is a story about competent output with fabricated provenance — which is a much harder problem, because quality inspection does not detect it.
Second: the deception was load-bearing. The AI was not incidental to the product. The absence of humans was the business model, and the claim of humans was the pricing power. Those are two different frauds stacked on each other, and only one of them is visible in the deliverable.
| Competent output with fabricated provenance is harder to catch than bad output. Quality inspection does not detect it — only provenance evidence does. |
AI Washing Due Diligence Has Been Running One Direction
The regulatory record is unambiguous about which way the enforcement lens has been pointed.
In March 2024 the SEC brought its first AI washing actions, charging investment advisers Delphia and Global Predictions with false and misleading statements about AI capabilities they did not have. The firms were censured and paid $225,000 and $175,000 respectively. Six months later the FTC launched Operation AI Comply, a sweep against companies using inflated AI claims to sell services. In April 2025 the Department of Justice took it criminal, indicting Albert Saniger, founder of the shopping app Nate, over claims that its “AI” checkout was actually hundreds of contractors in call centers in the Philippines and Romania. Prosecutors alleged the app’s real automation rate was effectively zero after roughly $40 million raised, describing it as a false narrative about innovation that never existed.
Every one of those cases punishes the same lie: we have AI and we don’t.
Research Gold is the mirror image: we have humans and we don’t. And the enforcement infrastructure for that direction is thinner, newer, and mostly not American.

Figure 2: US penalties for overstating AI have been modest. The EU ceiling for failing to disclose AI is two orders of magnitude larger. Sources: SEC Press Release 2024-36; Regulation (EU) 2024/1689, Arts. 50 and 99.
The Article 50 Clock Started Nine Days Ago
On August 2, 2026, Article 50 of the EU AI Act became applicable. It is the transparency layer of Regulation (EU) 2024/1689, and it is narrower and far more mechanical than the high-risk regime that dominated the compliance conversation. Its first obligation is the relevant one here: providers of AI systems that interact directly with people must ensure those people are informed they are dealing with an AI, unless it would be obvious to a reasonably observant person in that audience.
As Cooley notes in its client alert, the duties split between providers and deployers, with a limited extension to December 2, 2026 for marking content generated by systems already on the market. Disclosure buried in terms of service does not satisfy it; the information has to be perceivable in the interaction itself. The obligation reaches any provider whose product touches EU users, and it applies to systems already in production from day one. Penalties under the Act’s transparency tier run to €15 million or 3% of global annual turnover.
Put plainly: an AI sales agent that answers a prospect’s direct question about whether it is human by saying yes is not a gray area under Article 50. It is the paradigm case.
The counterweight is that US federal posture is moving the other way. The FTC reopened and set aside its Rytr order in December 2025, citing the current administration’s AI executive order and action plan. A deal team that assumes convergence between US and EU AI disclosure regimes is underwriting a jurisdictional divergence that is currently widening, not closing. Any target with meaningful EU revenue carries the Article 50 exposure regardless of what happens in Washington.
The Margin Arithmetic That Makes Human-Washing Rational
The economic logic here deserves attention because it is what turns a single bad actor into a category risk.
A systematic review done properly is not a light lift. Sebastian Rowan, a PhD candidate in civil and environmental engineering at the University of New Hampshire who first flagged Research Gold to 404 Media, described reading more than 250 articles end to end for a single meta-analysis. Protocol registration, dual screening, extraction, risk-of-bias appraisal, and synthesis is conventionally eighty to a hundred and twenty hours of credentialed analyst time.
At $75 to $150 an hour, that is $6,000 to $18,000 of cost against a $1,900 price. The human-delivered version of this product has deeply negative gross margin. The model-delivered version costs tens of dollars.

Figure 3: Illustrative model built from public rate ranges, not a vendor disclosure. Assumes 80–120 analyst hours at $75–$150/hr for human delivery; inference, tooling and overhead only for AI delivery. Source: DevelopmentCorporate.com estimate.
Figure 3 is directional, and it is the point: whenever a service is priced below the cost of the labor it claims to use, the pricing itself is the tell. That is not a moral observation. It is a diligence procedure.
This is the same structural pattern we documented in the Klarna Effect — a headline transformation narrative whose economics only worked because the human capital being removed was load-bearing in ways the model did not capture. It is also the pattern behind AI startup ARR quality problems more broadly: revenue that looks like software but is delivered like services, or services revenue reported at software margins because of an undisclosed substitution.
For an acquirer, the practical translation is direct. Anomalously high gross margin in a services-inclusive line item is now a provenance question, not a compliment. So is anomalously low pricing against a credentialed-labor promise. Both of them used to read as operating leverage. Now they read as an open question about what is actually behind the delivery.
What Provenance Failure Costs a Buyer After Close
The tail risk is not the discovery itself. It is the remediation across everything the compromised work touched.
Academic publishing already ran this experiment. When paper mills flooded journals with fabricated manuscripts, Hindawi retracted more than 8,000 articles and parent company Wiley closed nineteen journals, disclosing tens of millions in lost annual revenue. A bibliometric study in Scientometrics tracked 10,409 retracted paper-mill articles through 2024 and found retraction lag stretching to two and a half years — meaning the contaminated work sat in the literature, being cited, for years before removal.
That lag is the mechanism that should worry buyers. Provenance failures do not surface at the moment of delivery. They surface later, through a third party, after the output has been embedded in downstream decisions — the same compounding dynamic we mapped in our analysis of AI-hallucinated citations as a due diligence gap and of hallucinated consulting research contaminating enterprise intelligence.
There is a second liability here that is easier to price and easier to miss: identity misappropriation. Research Gold published real named professionals’ photographs and biographies without consent. Berrio said she was preparing a formal takedown request. In a target company, that same fact pattern — a customer-facing roster, credential page, or “meet the experts” section populated with people who never agreed to be there — is a right-of-publicity exposure, a potential defamation exposure, and a representation-and-warranty problem all at once. It sits in marketing, which is precisely where diligence tends not to look.
| Provenance failures do not surface at delivery. They surface later, through a third party, after the output is already embedded in downstream decisions. |
A Human-in-the-Loop Verification Protocol for Diligence Teams
Human-in-the-loop verification is not a governance abstraction. It is a set of evidentiary requests, and most targets cannot currently satisfy them. Our work on agentic AI governance as an unpriced liability and on enterprise AI security diligence makes the same point from adjacent angles: the question is never what the vendor asserts, it is what the vendor can evidence.
Seven AI washing due diligence questions to add to the request list for any target with a service, review, support, or expert-delivery component:
- Which delivery steps involve a human, named by role, and what artifact proves it? Timesheets, reviewer sign-off logs, and version histories are evidence. An org chart is not.
- Is every person on the public-facing site an employee or contracted party with documented consent to be listed? Ask for the consent records. This is a five-minute check that would have caught Research Gold instantly.
- What is the gross margin on each service line, and does it reconcile with the labor the marketing claims? Reconcile the pricing to the promised input cost. Divergence is the tell.
- Do customer-facing AI agents disclose that they are AI at the point of interaction? Not in the terms of service. In the conversation. Then test it: ask the agent directly and keep the transcript.
- What is the EU revenue exposure, and is there an Article 50 compliance assessment dated on or before August 2, 2026? Absence of a dated assessment on a target with EU users is a live regulatory finding, not a future risk.
- Have customer contracts made human-delivery representations? Search the contract corpus for “human,” “expert review,” “manual,” and “certified.” Each hit is a representation that must be independently substantiated.
- What happens to the deliverable if the provenance claim is wrong? For regulated outputs — clinical, legal, financial, safety — model the remediation cost across every downstream artifact, not just the refund.
What Counts as Human-in-the-Loop Evidence
The distinction that matters is between attestation and artifact. A management representation that humans review the output is an attestation. A sampled reviewer log, tied to specific deliverables, with named individuals whose employment can be confirmed, is an artifact. In a diligence process, only the second one survives contact with a dispute.
Buyers should sample. Pick five deliverables at random, trace each to the named human who produced or reviewed it, and confirm that person exists, was employed at the time, and recognizes the work. If a target cannot complete that exercise in a week, that answer is itself the finding — the same conclusion we reached about AI-assisted screening workflows that survive as claims but not as governed processes.
What AI Washing Due Diligence Means Across the Table
The same fact pattern produces three different action lists depending on which side of the transaction you sit on.
| FOR PE AND VC INVESTORSAdd provenance verification to the quality-of-earnings workstream, not the technology workstream. A services-inclusive revenue line delivered by undisclosed AI is simultaneously a margin-quality issue, a customer-contract representation issue, and a regulatory exposure — and it can flip from asset to liability the day a journalist or a customer runs the test that Maiberg ran.Where marketing has made an explicit human-delivery claim, that claim belongs in the representations and warranties with a specific indemnity, not folded into a general compliance rep. Note also that the asymmetry cuts the other way: a target that can evidence provenance has a differentiator the market is not yet pricing, consistent with the moat-quality framing in the a16z SaaS moat scorecard. |
| FOR SAAS FOUNDERS APPROACHING EXITAudit your own site before a buyer does. Every named person on your team page, advisory board, and customer-success roster should have documented consent. Every “expert-reviewed,” “human-verified,” or “certified by our team” claim in your marketing and your MSAs should have an artifact behind it.If you use AI in delivery — most companies now do, legitimately — say so explicitly and describe the human control points precisely. Undisclosed AI in delivery is a diligence bomb. Disclosed AI with documented human oversight is a margin story, and buyers are increasingly ready to underwrite it as one, in line with the AI monetization evidence we analyzed in the 2026 SaaS benchmarks. |
| FOR ENTERPRISE CTOS AND CPOSPut a provenance clause in the vendor contract template. Require disclosure of AI use in delivery, name the human control points, and reserve audit rights against them. For anything feeding a regulated decision, require that the vendor identify the human accountable for the output.Test your existing vendors the cheap way: call the support line and ask the agent whether it is a person. What comes back tells you both what the vendor is doing and how honest it is about doing it — a question that becomes more consequential as more workflows shift toward autonomous agents, as covered in our analysis of agentic AI and SaaS spending. |
The Takeaway: Audit Both Sides of the AI Claim
Research Gold is a small company selling a $1,900 product to academics, and it will probably be gone before this analysis is a month old. The lesson is not about Research Gold.
The lesson is that for two years, the entire market — regulators, acquirers, buyers — has been auditing one side of a two-sided claim. We built the machinery to catch companies that pretend to have AI. We built almost nothing to catch companies that pretend not to.
The margin incentive now runs firmly in the second direction. AI delivery is cheap, human delivery is expensive, and customers still pay a premium for the human label. That is a durable arbitrage, and it will not stay confined to marginal websites selling to graduate students. It will show up in your data room, in a services line with a gross margin nobody could quite explain, backed by a team page nobody thought to verify.
AI washing due diligence, done in both directions, costs a week of work. The failure costs the whole revenue line.
| WORK WITH DEVELOPMENTCORPORATEDevelopmentCorporate LLC advises enterprise SaaS founders, boards, and investors on M&A strategy, AI-era due diligence, and exit preparation. If you are underwriting a target with service-delivery components — or preparing your own company for a buyer who will run these tests — start a conversation with our team. |
