AI Startup Screening: How VCs Use ChatGPT to Filter Pre-Seed Deals
The first investor to read your pre-seed enterprise software pitch is probably not an investor. Here is how the algorithmic filter works — and where it breaks.
AI startup screening is now the first gate most pre-seed enterprise software founders will ever face — and almost none of them know it happened. The pitch deck you spent six weeks refining was very likely summarized by ChatGPT or Claude in under a minute, compressed into three sentences, scored against a rubric, and routed to a pass pile before a human partner ever opened it. According to Affinity’s 2026 survey of nearly 300 private capital dealmakers, 85% now use AI to automate daily tasks, up from 76% a year earlier, and 82% of firms use AI for deal sourcing research.
The venture industry’s story about this shift is a story about efficiency: more decks processed, faster responses, less associate burnout. That story is true and it is also beside the point. The real question — the one that matters to the LPs funding these firms, the founders being filtered by them, and the enterprise buyers who will eventually depend on the companies that survive the filter — is what the algorithmic screen actually measures. The evidence emerging from controlled studies of large language models in investment contexts suggests an uncomfortable answer: LLM screening measures resemblance to historical patterns, at precisely the investment stage where the entire financial return depends on finding companies that do not resemble historical patterns.
That is the gap this analysis examines. Not whether VCs use generative AI to screen pre-seed enterprise software candidates — they do, pervasively — but what the screening layer systematically catches, what it systematically misses, and what both sides of the table should do about it.
The Adoption Sprint Nobody Audited
The speed of AI adoption in venture screening has no precedent in the industry’s own tooling history. CRM adoption took a decade. Data-driven sourcing platforms took most of another. Generative AI screening went from novelty to default in roughly three years. Compiled industry statistics show the share of VC firms using AI tools for deal sourcing rising from 22% in 2021 to 78% by 2024, with solo GPs reporting that AI now handles roughly 75% of their initial screening. By 2026, 90% of VCs expect to use AI across all sourcing activity.

Figure 1: AI adoption in VC deal screening, 2021–2026. Sources: Gitnux statistics compilation; Affinity dealmaker surveys; DevelopmentCorporate.com analysis.
But adoption headlines conceal a production gap that should look familiar to readers of our work on AI hallucination rates as a due diligence crisis: the difference between what a tool does in a survey response and what it does in a governed workflow. Capitaly’s analysis of institutional VC AI adoption, based on conversations with operators at more than 20 institutional funds, found that fewer than 12% of institutional VC funds have functional AI-assisted deck triage workflows actually in production. The rest are, in Capitaly’s phrase, in pilot purgatory — running prompts through Claude or ChatGPT in private Slack channels while humans still read every deck.

Figure 2: The screening gap — claimed AI usage vs. governed production workflows. Sources: Affinity (2026); Capitaly; DevelopmentCorporate.com.
This 73-point gap is the single most important fact about AI startup screening in 2026, because it defines the risk profile. A governed screening system has audit logs, calibration checks, and someone accountable for false negatives. A ChatGPT prompt in a private Slack channel has none of those things — but it produces the same output: a recommendation about which founders get a meeting. The industry has deployed algorithmic gatekeeping at survey-response speed and governance at pilot-purgatory speed. The difference between those two speeds is where the errors live.
How AI Startup Screening Actually Works at Pre-Seed
Strip away the vendor marketing and the actual workflows in use across funds sort into three tiers, plus a proprietary end of the spectrum that operates on entirely different economics.
Tier 1: Inbox triage with general-purpose LLMs
This is the dominant mode, and it is exactly as unglamorous as it sounds. An associate — or at smaller funds, the GP — drops an inbound deck into ChatGPT or Claude with a standing prompt: extract the core claim, the ask, the founders’ backgrounds, and any prior exits; flag stage mismatch, sector mismatch, or obvious red flags; produce a two-to-three sentence summary. Per Capitaly’s reporting, brand-name funds including Sequoia, Accel, and Andreessen Horowitz run private LLM instances or Claude API integrations to auto-tag decks for sector, stage, and red flags — with the important caveat that at these firms the AI does markup, not decision-making. At the long tail of smaller funds and solo GPs, that caveat quietly disappears. When AI handles 75% of a solo GP’s initial screening, the summary is the decision for most of the inbox.
Tier 2: Extraction, enrichment, and benchmarking
The second tier moves from summarization to structured extraction. Investor-side pitch deck analysis workflows pull claimed metrics out of decks — team size, pilot counts, market sizing, burn — and compare them against portfolio benchmarks and sector norms. For a pre-seed enterprise software company, this is where claimed TAM meets an automated sanity check, and where a founder’s “$40B market” slide gets cross-examined by a Perplexity search before any human considers it. Firms use ChatGPT for summarizing decks and analyzing market trends, and Perplexity for citation-backed validation of pitch deck claims — a division of labor that has become standard enough to be boring.
Tier 3: Validation and background synthesis
The third tier is the background layer: automated founder research combining LLM search tools with LinkedIn, GitHub, and press coverage. At pre-seed enterprise software — where there is rarely meaningful revenue — founder background is the densest signal available, so this tier carries disproportionate weight. It is also the tier most dependent on what the models were trained on and what they can retrieve, a dependency we mapped in detail in our analysis of what ChatGPT, Claude, Gemini, Grok, and Perplexity are actually trained on. A founder whose track record lives in gated content, paywalled press, or thin public footprints is — to this tier of screening — a founder with no track record.
The proprietary end: Motherbrain and its cousins
At the opposite end of the spectrum sit purpose-built sourcing platforms that predate the ChatGPT era and now incorporate it. EQT’s Motherbrain, in operation since 2016, ingests headcount growth, web traffic, founder social activity, hiring composition, and news signals across millions of companies, scoring them continuously against the firm’s thesis. VentureBeat’s reporting documented that the platform directly identified multiple EQT investments — including Peakon — and enabled the firm to approach startups proactively, months before they raised. SignalFire’s Beacon platform has tracked data across millions of startups since 2013. These systems matter to the pre-seed screening story for one reason: they demonstrate what disciplined, purpose-built algorithmic sourcing looks like, and by contrast, how far a ChatGPT prompt in a Slack channel falls short of it. Notably, a 2024 SSRN study cited in Crustdata’s analysis of proprietary VC pipelines found that funds adopting ML-based sourcing became significantly more likely to invest outside traditional startup hubs — and that those outside-hub investments were more likely to IPO or reach unicorn status. Purpose-built systems, properly deployed, widen the aperture. The question is whether general-purpose LLM triage does the same — or the opposite.
What the screen structurally cannot see
Before examining what LLM screening gets wrong, be precise about what it cannot evaluate at all. V7’s guide for investor-side AI deck analysis — written by practitioners building these exact workflows, and therefore inclined to oversell them — is unusually candid on this point: nothing AI extracts from a deck tells you whether the founder can execute. How a founder responds to pushback, whether they have the self-awareness to pivot when wrong, whether they are someone a board can work with — none of it is in a slide deck, and therefore none of it is in the model’s input. Neither is relationship data. The question that most often decides a first-pass verdict at a well-networked fund — who do we know who knows this team? — lives in the partnership’s network graph, not in the document. And the most valuable catch an experienced sector investor makes at screening is spotting the flaw in a market-size claim that requires knowing what is not in the deck: the incumbent contract structure, the procurement cycle, the regulatory landmine the founder has not yet encountered.
At later stages, these blind spots are cushioned by data the model can see. At pre-seed enterprise software, they are most of the investable signal. The stage where AI screening has been adopted most aggressively — because inbound volume is highest and staffing thinnest — is the stage where the machine-readable share of the decision is smallest. That inversion is not a transitional artifact that better models will fix. It is a structural property of what pre-seed evidence is: private, relational, and behavioral, encountered in rooms rather than encoded in documents. Venture capital’s historical excess returns came from exactly this information asymmetry — knowing something the market does not — and a screening layer built entirely on public and founder-supplied text is, by definition, working from the information everyone has.
| For PE/VC InvestorsThe 12% production figure is a diligence question for your own fund, not just your targets. If your screening runs on ungoverned LLM prompts, you cannot answer the LP question that is coming: “What is your false negative rate, and how do you know?” Funds that can produce calibration data on their AI screen — decks passed, decks killed, outcomes tracked — will have a defensible answer. Funds that cannot are running an unaudited model with fund-returner-sized error bars. |
The Pre-Seed Problem: Screening Companies That Have No Data
Everything above applies to venture screening generally. Pre-seed enterprise software is the special case where the method is weakest, for a structural reason: at pre-seed, the evidence that LLM screening is good at processing does not exist yet.
A Series B screen has cohort retention, NRR, pipeline conversion, and burn multiples — quantitative inputs where automated extraction and benchmarking genuinely outperform manual review. A pre-seed enterprise software company has a deck, two founders, possibly a design partner or an unpaid pilot, and a narrative. As we documented in our 2025 win/loss reality check for pre-seed and seed enterprise SaaS, these companies are selling into conservative buyers with no brand recognition and a sales motion still being built in real time. There is no metrics layer to extract. What remains for the model to evaluate is the story and the pattern: does this founder, this market framing, this wedge resemble things that worked?
That is not a neutral question. LLMs are, by construction, consensus machines — they compress the distribution of their training data and retrieval corpus into the most probable continuation. Asked to evaluate a pre-seed narrative, the model rewards decks that resemble the median successful pitch as represented in its training data: recognizable framing, familiar market categories, founder biographies that rhyme with past winners. The pre-seed asset class, meanwhile, is a power-law business in which essentially all of the return comes from outliers — companies that at the screening moment looked wrong to most sophisticated observers. The screening tool is optimized for central tendency. The asset class is defined by tails.
Human screeners have their own well-documented pattern-matching pathologies, and it would be revisionist to pretend the pre-AI status quo was a meritocracy of ideas. But human pattern-matching at least varies — across partners, moods, networks, and theses — and that variance is itself a diversification mechanism for the ecosystem. When a growing share of funds runs first-pass screening through the same two or three foundation models, with similar prompts, the industry converges on a correlated filter. Correlated filters produce correlated blind spots. In public markets we would call this crowding. At pre-seed, nobody is even measuring it.
| For SaaS Founders Raising Pre-SeedAssume your deck’s first reader is a model. That means three practical changes. First, structure for extraction: state the core claim, the wedge, the team’s relevant history, and the ask in plain declarative text — not in imagery, not in a voiceover you plan to deliver live. Second, make your public footprint machine-verifiable: when Tier 3 screening searches your name and company, what it retrieves is your track record. Third, expect your market-size claims to be checked by a search engine in real time — a defensible bottom-up number beats an impressive top-down one. |
The Gap Thesis: What LLM Screening Measures vs. What Predicts Outcomes
Readers of this site know the recurring pattern in our AI coverage: the benchmark number, measured under favorable conditions, diverges from production reality — and the gap is the story. Vendor hallucination claims of 0.9% versus observed rates of 69–88% on complex legal tasks. Benchmark agent performance versus the autonomy gap in deployed agentic systems. AI screening of pre-seed deals is the same structure, and for once we have controlled evidence.
The models systematically underfund
A 2026 study in Springer’s Digital Finance journal, “Algorithmic personalities and the myth of neutrality,” ran a controlled simulation directly on point: GPT-4o, Claude 3.5 Sonnet, and DeepSeek-V2 each evaluated 20 real startup pitch decks across industries and funding stages, with each model pair evaluated five times under identical conditions to separate noise from persistent behavior. Three findings should reorganize how every LP and GP thinks about LLM screening. First, none of the models replicated human funding decisions with sufficient accuracy for autonomous use — the authors’ words, not ours. Second, the models showed a consistent tendency to underfund relative to observed market outcomes; in Claude’s case, underfunding appeared in 79% of evaluations. Third, the differences between models were systematic and statistically significant — funding recommendations, evaluation scores, and expressed confidence all varied by model in reproducible ways.
Model choice is an investment decision nobody made deliberately
Sit with that third finding. If GPT-4o and Claude produce systematically different verdicts on identical decks, then a fund’s choice of screening model — typically made by whichever associate set up the workflow, on the basis of a personal subscription — is silently embedding an evaluative personality into the top of the funnel. The study’s authors describe the industry’s working assumption, that sufficiently advanced models produce broadly comparable outputs, and then find limited support for it. No investment committee voted on this. Most cannot tell you which model screens their inbox, let alone how its biases differ from the alternative’s.
A second controlled experiment reinforces the pattern from a different angle. A retrospective study, also in Digital Finance, gave three LLM configurations — GPT-4, Claude Opus, and a consensus ensemble — the identical information available to Berkshire Hathaway at 21 decision points across 2022–2024. The models demonstrated sophisticated financial reasoning and tighter risk management, yet returned 4.72–18.01% cumulatively against Berkshire’s 42.12%, dragged down by systematic behavioral tilts: excessive momentum bias, chronically low market exposure, and premature profit-taking. Translate that to screening: given the same inputs as an elite human investor, the models did not make random errors — they made directional ones, consistently favoring the recent, the familiar, and the already-validated. Momentum bias in public equities is underperformance. Momentum bias at pre-seed — favoring the category that just got hot over the one about to — is the exact opposite of the job.
Arbitrary variance dressed as judgment
The evaluation pathologies documented in adjacent domains transfer directly. MIT Computational Law’s study of LLM resume screening found that across more than 2,000 observations, ChatGPT overwhelmingly selected the first resume presented when candidates were equally qualified — and when explicitly instructed to avoid position bias, the bias simply moved to a different position, with several positions never selected at all. A pre-seed screen that batches ten decks into a context window inherits exactly this failure mode: ordering effects masquerading as evaluation. And all of this sits on top of the fabrication risk we quantified in our due diligence analysis of hallucination rates — models confidently asserting founder histories, competitor landscapes, and market figures that do not survive verification.
The false-negative asymmetry
Here is why this matters more at pre-seed than anywhere else in the capital stack. In most screening contexts — hiring, credit, procurement — false positives and false negatives carry comparable costs, so a filter that trades a few missed gems for large efficiency gains is rational. Venture economics are the opposite. The cost of a false positive is one wasted meeting. The cost of a false negative is, occasionally, the fund. A screening layer that is systematically conservative — that underfunds relative to market outcomes in three-quarters of its evaluations, that penalizes unfamiliar framing, that reproduces the consensus of its training corpus — is a machine for manufacturing false negatives in the one asset class where false negatives are catastrophic and invisible. Invisible is the operative word: a passed deck that becomes a unicorn generates no error report. The feedback loop that would correct the filter does not exist at the timescale on which the filter operates.
Founders Are Optimizing for the Algorithm — Which Contaminates the Signal
Every filter trains its population. Within roughly a year of AI screening becoming ambient, founder behavior adapted: decks structured for LLM extraction, narrative language tuned to score well on rubric prompts, and — the sophisticated version — deliberate cultivation of the public footprint that Tier 3 screening retrieves. This is generative engine optimization applied to fundraising, and it is the fundraising mirror of the dynamic we documented on the demand side, where 94% of enterprise B2B buyers now use LLMs to research software purchases and vendors invisible to the models are eliminated from deals they never knew they were in.
The strategic implication cuts both ways. For founders, a citation footprint — press coverage, structured public content, verifiable founder history — is now a fundraising asset with measurable screening consequences, exactly as it has become a sales asset and, as we have argued, an M&A valuation asset. For investors, founder GEO is signal contamination: the decks that score best under LLM screening are increasingly the decks engineered to score best, which are not necessarily the companies most likely to compound. The screen selects for founders skilled at satisfying the screen. Anyone who has watched SEO eat organic search quality over two decades knows how this movie ends — the filter degrades as the population optimizes against it, and the arms race consumes the signal it was built to detect. The same adversarial dynamic corrupting synthetic and AI-mediated market research is now running inside venture deal flow.
There is a subtler consequence for how screening scores should be read. In a pre-optimization environment, a high LLM screening score was weak positive evidence — a legible deck from a legible founder. In an adversarial environment, the same score carries a second interpretation: a founder who invested in scoring well. Distinguishing the two requires exactly the verification the screen was supposed to eliminate. Meanwhile the inverse signal quietly appreciates. A deck that scores poorly because its framing matches no existing category, from a founder whose footprint is thin because they spent the last four years inside an enterprise fixing the problem they are now productizing, is precisely the profile the anti-portfolio literature says funds regret missing. The uncomfortable arithmetic: as founder-side optimization spreads, the expected value of a low screening score rises relative to a high one — not because bad decks became good, but because the high scores became crowded and the low scores became where the unpriced deals live.
| For Enterprise CTOs and CPOsYour vendor shortlist is downstream of this filter. The pre-seed enterprise software companies that survive AI screening — and therefore exist as fundable vendors three years from now — skew toward familiar categories, legible narratives, and strong public footprints. Genuinely novel architectural approaches are disproportionately filtered out before your procurement process ever sees them. If your innovation sourcing depends entirely on the venture-backed pipeline, you have inherited the biases of the models screening it. Direct scouting, university spinouts, and open-source traction signals are now diversification tools against a correlated upstream filter. |
A Working Framework for AI-Screened Deal Flow
None of this argues for abandoning AI screening — the volume economics are irreversible, and tools like Motherbrain prove algorithmic sourcing can widen rather than narrow the aperture when built deliberately. It argues for treating the screening layer as a model with failure modes, and managing it the way our SaaS moat scorecard treats competitive claims: verify structurally, not rhetorically. Five questions operationalize this.
- Which model, which prompt, which version? If a fund cannot name the model and prompt template screening its inbox, its screen is unaudited. The Springer findings make model identity a material evaluative choice, not an implementation detail.
- What is the escape hatch for pattern-breaking deals? The correct design routes low-scoring-but-anomalous decks to human review rather than auto-pass. Conservatism plus anomaly detection beats conservatism alone in a power-law asset class.
- Is anything verified before it influences a decision? LLM-generated founder background and market claims must be treated as leads to verify, not facts to rely on. Hallucinated screening inputs are hallucinated investment theses.
- Are outcomes tracked against screening verdicts? The only honest calibration is longitudinal: log every pass, revisit annually, count the anti-portfolio. A fund that has never measured its screen’s false-negative rate does not have a screening system; it has a habit.
- Is ordering randomized and are batches controlled? Position bias is a documented, reproducible failure mode. Screening decks in fixed batch order imports it wholesale.
For LPs, these five questions convert directly into diligence language for fund commitments. The 2026 vintage is the first in which a meaningful share of a fund’s anti-portfolio will have been shaped by a foundation model’s evaluative personality — and the first in which “describe your AI screening governance” is a question with a checkable answer. A GP who responds with tooling names is describing adoption. A GP who responds with calibration data, override rates, and an anomaly-routing design is describing a system. The difference will not show up in this year’s deployment pace. It will show up in DPI a decade from now, in the deals the second GP saw and the first one never knew existed.
The Bottom Line
VCs use generative AI to screen pre-seed enterprise software candidates because the volume math makes it unavoidable — and, deployed with governance, it is a genuine edge. But the evidence now on the table is specific: the models systematically underfund relative to market outcomes, disagree with each other in reproducible ways, exhibit arbitrary positional biases, and reward resemblance to consensus in an asset class whose returns come exclusively from deviation. Meanwhile, fewer than one fund in eight runs this machinery inside a governed workflow, and the founder population is actively optimizing against the filter.
The gap between AI screening as surveyed and AI screening as governed is the same gap we keep finding across the enterprise AI landscape — between benchmark and production, between demo and deployment, between the number in the deck and the number in the field. At pre-seed, that gap has a special property: its costs are invisible, deferred, and occasionally fund-sized. The firms that win the next vintage will not be the ones that screened fastest. They will be the ones that knew what their screen could not see.
DevelopmentCorporate LLC advises enterprise SaaS companies and investors on exit strategy, acquisition diligence, and valuation positioning — including AI-readiness assessment of both deal processes and screening pipelines. If you are an LP evaluating a fund’s AI screening claims, or a founder positioning for an algorithmically screened raise, contact us for a confidential discussion at developmentcorporate.com.
