A laptop displaying AI prompt discovery text next to chat log transcripts, a magnifying glass inspecting prompts, and litigation files in a corporate boardroom.
| |

AI Prompt Discovery: The M&A Risk Hiding in Your Chat Logs

A 3M expert witness handed over 350 pages of ChatGPT prompts. The report was defensible. The record was not. That distinction is about to reprice AI governance in enterprise SaaS transactions.

AI prompt discovery arrived in American courtrooms this month, and it did not arrive the way the governance literature predicted. For two years, the consensus risk framework around enterprise AI has centered on output quality: hallucinated citations, fabricated data, and unverified claims. We have covered that risk extensively — AI hallucination rates remain the most underpriced risk in enterprise SaaS transactions. Vendors sell verification layers. Deal teams ask about accuracy benchmarks. Disclosure schedules address reliability.

Then a court in Harris County, Texas obtained an expert witness’s ChatGPT transcripts, and the failure mode turned out to be somewhere else entirely.

The output was not the problem. The record was.

What Actually Happened in the 3M Expert Witness Case

The underlying litigation concerns the January 24, 2020 explosion at Watson Grinding and Manufacturing in northwest Houston. A 2,000-gallon propylene tank leaked overnight through a degraded, poorly crimped rubber welding hose. The facility’s gas detection system never sounded an alarm. When a worker flipped an ordinary light switch the next morning, the accumulated gas ignited. Three people died. Dozens were injured. Hundreds of nearby homes and businesses were damaged or destroyed.

The U.S. Chemical Safety and Hazard Investigation Board attributed the incident to a combination of the degraded hose, an unclosed manual shutoff valve, and non-functioning safety systems. Plaintiffs sued Watson Grinding and 3M, which had been contracted to inspect and service the gas detection system. More than 2,600 plaintiffs were consolidated into multidistrict litigation. Reporting from 404 Media established what happened next.

3M retained Josh Autenrieth of KnightHawk Engineering as an expert witness. Plaintiffs’ counsel suspected AI authorship during discovery and moved to obtain his ChatGPT logs. The court agreed. 3M’s team produced roughly 350 pages of transcripts — which, per the reporting, also contained publicly accessible share links to the full conversations.

The prompts were not ambiguous. Autenrieth asked the model to create an expert witness report defending the standard of care at 3M, and to show how 3M was zero percent at fault for the explosion. He uploaded hundreds of court records and asked the model to construct a defense assigning responsibility to Watson Grinding. At one point he uploaded a photograph of a gas detector — the device at the center of the entire case — and asked the model what he was looking at. Per Above the Law, he was retained for upwards of $90,000.

“He acknowledged the prompts he put in were biased toward 3M to help 3M win the case.” — Will Moye, plaintiffs’ counsel

Here is the detail that matters most for anyone underwriting AI risk. When Autenrieth asked the model to critique its own draft, it flagged the “zero percent responsible” framing as an easy target for opposing counsel. He removed the line. The verification loop worked. The final deliverable was cleaner than the first draft.

It did not matter. The report was never what got tested. The prompts were.

The Gap Thesis: Verification Fixes the Output, Not the Record

Every AI governance framework currently sold into the enterprise is built to improve the artifact. Citation checkers validate references. Retrieval grounding constrains claims. Human-in-the-loop review catches errors before delivery. These controls are real and they work — on the dimension they measure.

None of them touch the process record. And in litigation, the process record is the more dangerous document, because it reveals intent rather than error. An inaccurate report is a competence problem, and competence problems can be rehabilitated on redirect. A prompt that specifies the conclusion before the analysis exists is a credibility problem, and credibility problems are terminal.

This is the gap. Organizations have invested in making AI output defensible while generating, in parallel, a complete and admissible transcript of how that output was steered. The better the verification layer, the wider the divergence between what the deliverable says and what the record shows. Our analysis of AI hallucinations as a ticking M&A time bomb in legal practice framed the question as whether a target’s AI can withstand scrutiny if challenged. The 3M deposition answers a harder version of it: whether the operator can.

Figure 1: Watson Grinding MDL outcomes against 3M — three juries, three results.

Figure 1 shows why expert credibility carries real financial weight in this specific litigation. The same defendant, facing substantially the same factual record, has drawn radically different outcomes. In November 2025 a Harris County jury awarded $118 million to seven plaintiffs and assigned 3M forty-nine percent of the fault. In May 2026 a different jury cleared 3M entirely, finding Watson Grinding solely negligent and awarding roughly $1.9 million against other parties. In August 2026 a third jury returned $61.5 million against 3M for twenty-four survivors.

The variance between those outcomes is not explained by the physical facts, which did not change. It is explained by what each jury found persuasive — which is precisely where expert testimony operates. A discredited expert is not a procedural footnote in a case with this spread. It is a nine-figure variable.

How AI Prompt Discovery Became Settled Practice

The 3M deposition is the most vivid example, but it is not the origin of the doctrine. Courts have been building toward it for two years, as we documented in our coverage of escalating judicial sanctions for AI-generated filings. The legal groundwork for prompt production specifically was laid over the preceding fifteen months, and it was laid quietly.

Figure 3: The precedent chain that made AI prompt discovery routine.

In May 2025, the U.S. District Court for the Southern District of New York ordered OpenAI to preserve and segregate ChatGPT output logs that would otherwise have been destroyed under a thirty-day deletion policy. That order was lifted in October 2025 after challenge, but the reasoning survived. On January 5, 2026, District Judge Sidney H. Stein affirmed two discovery orders requiring OpenAI to produce a sample of twenty million de-identified user conversations in the consolidated copyright litigation brought by news organizations. As Robinson+Cole noted in its client analysis, the provider retains tens of billions of such logs in the ordinary course of business.

Three features of that ruling deserve attention from anyone building an AI governance program.

  • The court treated conversational AI data — prompts and outputs together — as an ordinary discoverable record category. No new privilege was recognized.
  • User expectations of privacy were considered and found insufficient to block civil discovery, on the reasoning that users voluntarily submitted the communications.
  • Enterprise and zero-retention tiers were carved out. The exposure fell on standard consumer accounts — the tiers most individual professionals actually use.

That third point is the one most organizations have not internalized. The distinction between a consumer subscription and an administered enterprise deployment is no longer a procurement preference. It is a determination of who controls the evidentiary record.

Deployment Tier Is Now a Legal Control, Not a Purchasing Decision

Most enterprises made their AI tier decisions in 2023 and 2024, on cost and feature grounds, before any of this precedent existed. Many never made a decision at all: employees expensed consumer subscriptions or used personal accounts, and the organization inherited the resulting exposure without ever seeing a contract. The McKinsey Lilli breach made the same structural point from the security side — an AI platform ingesting sensitive workflows at scale is a governance-critical asset, not an IT tool.

Figure 2: Discovery exposure varies sharply by deployment tier. Illustrative model.

Figure 2 maps four exposure dimensions against four common deployment configurations. The pattern is consistent: consumer tiers concentrate risk on every axis, while administered enterprise and zero-retention deployments compress it. Note the share-link dimension specifically. In the 3M matter, the transcripts were reportedly accompanied by publicly accessible conversation links — a convenience feature that converted a private working session into published material.

For enterprise buyers, this reframes shadow AI. The conventional concern about unsanctioned AI tools is data leakage: proprietary information leaving the perimeter. The discovery concern runs the other direction. Shadow AI creates a body of evidence about your organization’s reasoning that sits in a third party’s infrastructure, outside your legal hold apparatus, retrievable by anyone who can articulate relevance to a magistrate judge.

Four Ways the Prompt Record Turns Against You

The 3M matter illustrates one failure mode vividly, but it is not the only one. In practice the prompt record creates exposure through four distinct mechanisms, and they call for different controls.

1. Conclusion-first prompting

The operator specifies the desired finding and asks the model to construct support for it. This is the 3M pattern. It is the most damaging because it establishes bias directly, in the operator’s own words, with a timestamp. No amount of downstream editing cures it. The control is training, not tooling: professionals producing work product that may be tested need to understand that the prompt is a draft of the work, discoverable on the same terms.

2. Scope drift between prompt and deliverable

The operator asks the model to perform an analysis it is not equipped to perform, then presents the result as though it rested on independent expertise. Uploading a photograph of the central piece of equipment in a case and asking what it depicts is scope drift in its purest form. The deliverable claims twenty years of domain experience; the record shows the expertise was outsourced mid-task.

3. Unreproduced reliance

The operator accepts a model output without independently verifying the underlying reasoning, then signs a document attesting to it. The output may even be correct. The problem surfaces when opposing counsel asks the operator to explain the derivation and the record shows there was none to explain. This is the inverse of the pattern we described in AI washing due diligence and the human-in-the-loop gap: there, a human process was claimed and not performed. Here, human judgment is attested and not exercised. In a transaction context, this is the mechanism most likely to reach a representation in the purchase agreement.

4. Uncontrolled publication

The operator uses a share link, a public workspace, or a personal account, and the conversation becomes accessible without any legal process at all. This is the cheapest failure to prevent and the one most frequently overlooked, because it is a product feature rather than a policy violation. It is also the only one of the four that can be remediated by configuration alone.

Categories one and two are behavioral. Category three is procedural. Category four is technical. An AI governance program that addresses only the technical layer — which describes most programs currently in production — has closed one quarter of the exposure and can reasonably believe it has closed all of it.

What AI Prompt Discovery Means Across the Table

The same fact pattern produces three different action lists depending on where you sit.

For PE and VC InvestorsAdd prompt retention and deployment tier to the technology diligence workstream. A target running material workflows on consumer AI accounts has an uncontrolled evidentiary surface that transfers with the asset.Ask whether AI-assisted work product supported any representation in the transaction documents. Financial models, market sizing, security questionnaire responses, and quality-of-earnings support are all candidates.Treat the prompt record as a distinct indemnity question. Existing AI reps generally address output accuracy and IP contamination. They rarely address whether a discoverable process record contradicts a disclosed conclusion.Raise it with R&W underwriters before they raise it with you. Underwriters already ask about hallucination exposure; prompt discoverability is the same conversation one layer down.
For SaaS Founders Approaching ExitYour diligence materials are being produced through tools that log. A prompt asking the model to strengthen the net revenue retention narrative is the sell-side equivalent of asking it to show zero percent fault.Move exit-process AI work onto an administered tier now, before the data room opens. Retroactive migration does not retire the logs already created on consumer accounts.Audit for public share links across the deal team. This is a ten-minute exercise with an asymmetric payoff.Document your AI governance posture as a diligence asset rather than a diligence liability. Buyers reward the seller who arrives with the answer already prepared.
For Enterprise CTOs and CPOsInventory AI interaction data as a record category alongside email and messaging. If it is not in the retention schedule, it is not under legal hold.Make tier selection a governance decision with legal sign-off, not a line item in the software budget.Disable or restrict public conversation sharing at the tenant level where the platform permits it.Update litigation hold notices and custodian interview templates to name AI tools explicitly. Custodians will not volunteer chat histories they do not think of as records.

An AI Prompt Discovery Diligence Framework

The following seven questions belong in the technology and legal workstreams of any enterprise SaaS transaction where AI is materially embedded in the product or the operating model. As our M&A due diligence checklist notes, legal and regulatory review is the workstream most likely to surface deal-breaking liabilities that financial analysis misses. Prompt discoverability now belongs in that column.

  1. Which AI platforms are in use across the organization, on which tiers, and under whose contract? Include personal and expensed accounts. The answer a target gives here is frequently incomplete, and the incompleteness is itself the finding.
  2. What is the provider-side retention posture for each platform? Distinguish between what the user interface promises about deletion and what the provider contractually retains.
  3. Are AI interaction logs included in the document retention schedule and reachable by a litigation hold? If the answer is no, the organization cannot preserve what it cannot see.
  4. Has public conversation sharing been used, and can it be audited? Any published link is discoverable by anyone, without a subpoena.
  5. Which representations in the transaction documents rest on AI-assisted work product that was not independently reproduced? This is the direct analogue of the expert witness problem.
  6. Does the target sell AI features that log customer prompts? If so, the target is a custodian of its customers’ discoverable records, which is a contractual and reputational exposure that belongs in the model.
  7. Is there any documented instance of a prompt specifying a desired conclusion in a compliance, safety, financial, or regulatory context? One such instance changes the risk profile of the entire category.
An inaccurate output is a competence problem. A prompt that specifies the conclusion is a credibility problem. Only one of those is survivable.

Why This Reprices AI Governance

There is a straightforward commercial reading of all this, and it is not that enterprises should use less AI.

The governance, risk, and compliance software category has been absorbing institutional capital for two years on the strength of cybersecurity and AI reliability mandates. AI prompt discovery extends that thesis into adjacent territory: prompt logging and retention infrastructure, AI-aware e-discovery tooling, tenant-level policy enforcement, and audit trails that distinguish assistive use from wholesale delegation. Our analysis of federal court AI adoption data identified the same market gap — judicial scrutiny is not slowing AI adoption, it is filtering it toward tools with verification and audit architecture built in. Every general counsel who reads about the 3M deposition acquires a procurement requirement.

For acquirers, the asymmetry is favorable. Prompt discovery exposure is cheap to diagnose and expensive to remediate after close. A handful of questions in diligence surfaces it. Discovering it during a post-close dispute, when the logs are already in a plaintiff’s hands, does not.

For sellers, the calculation is the reverse and just as clear. This is a controllable exposure with a short remediation runway. Tier migration, share-link audits, and retention schedule updates are weeks of work, not quarters. Companies that complete that work before the data room opens will price the risk at zero. Companies that do not will negotiate it at the buyer’s valuation.

The Deliverable Is No Longer the Only Document

Josh Autenrieth produced a thirty-page expert report that had been reviewed, critiqued, and revised. By the standards of every AI verification framework on the market, the process worked. The weakest claim was identified and removed before submission.

What ended his testimony was not in the report. It was in the record behind it — a record he did not consider a record, generated on a platform he did not control, produced under an order he did not anticipate.

That is the shape of the exposure. Not the output. The record. And in a market where AI capability commands a measurable valuation premium, the gap between what a company’s AI produces and what its AI logs reveal is a diligence question that almost nobody is asking yet. It belongs in the same category as the other AI risks standard deal rooms have no process to measure.

The organizations that ask it first will be the ones setting the price.

DevelopmentCorporate LLC advises enterprise SaaS founders, private equity investors, and strategic acquirers on M&A transactions, exit strategy, and AI-era due diligence frameworks. If you are evaluating an acquisition target with material AI exposure, or preparing your company for a diligence process, contact us to discuss how prompt discoverability fits into your risk framework.

Similar Posts