Ahammad Shibilbiology · capital · writing
Writing / Atoms & Cells

biology · 19 min read

When Biology Becomes Information

The fifty-year compression curve, the hundredfold fall in platform value it has already produced, and the one architecture that survives it. A field guide for the people building biotech platforms, the people funding them, and the people allocating to the funds that do.

In 1980, Genentech went public at $35 a share. It opened at $88. The company had, at that point, no approved product — what it had was a technique. A small cohort of people knew how to splice a human gene into a bacterium, coax the bacterium to express a folded protein, and purify that protein to clinical grade. The underlying science was not secret; Cohen and Boyer had published the recombinant-DNA method in 1973 and Stanford licensed it widely. What was scarce was the execution. Knowing how to scale a fermenter and purify a protein at purity was a craft held by a few dozen people in a few dozen institutions, and that craft was worth a decade of monopoly. Genentech, Amgen, and the cohort around them built it into companies that are today worth, in aggregate, hundreds of billions of dollars.

In May 2024, DeepMind published AlphaFold 3 in Nature — a model that predicts the three-dimensional structure of proteins and their interactions with startling accuracy. It was one of the most significant pieces of biology released that decade. Within roughly two months, an open-source rival shipped comparable structural-prediction performance, and Meta released a stack in the same class. The lead time the pioneer held was not a decade. It was a quarter.

That is the whole story of biotech in two data points. The lead time a platform builder can hold over the rest of the field has fallen from ten years to ten weeks across forty years, and it has fallen for one reason: biology is becoming an information technology, and information does not stay scarce. Everything else in this essay is the working-out of that single sentence — what it has already done to where value accrues, what it is doing right now, and what it means for you depending on whether you are building a biotech platform, funding one, or allocating to the people who do.

The one idea

A biotech platform is, at bottom, a bet on scarcity. You build something — a technique, an instrument, a dataset, a model — that lets you produce drug candidates faster or better than anyone else, and you capitalize the gap between what you can do and what the field can do. The size of the prize is the size of the gap, multiplied by how long you can hold it.

The trouble is that the gap is closing, and it is closing faster every generation, because the part of biology that any given platform monopolizes keeps turning into information. A technique held in a few people's hands is scarce. The same technique, once it has been written down as a protocol, embedded in a model, packaged into an open-source tool, and run on rented cloud compute by a graduate student in any city with payment rails — is not scarce at all. The frontier of craft — the part of drug discovery that is still expensive, tacit, and geographically concentrated — is where the platform premium lives. And that frontier has been retreating for fifty years.

This is why the same process keeps repeating, and why it keeps accelerating. A closed-source pioneer opens a gap. An open-source rival closes it. The value migrates to the next frontier of craft. Each time the cycle runs, it runs faster, because the machinery that converts craft into information — published models, cloud infrastructure, mobile talent, composable tools — gets more powerful. Run that loop five times across fifty years and you get a curve. Run it forward and you get the present moment, which is the most interesting one yet, because the layer of biology now turning into information is discovery itself.

Hold that idea. Everything below is the evidence.

The compression, told as a story

The cleanest way to read the last fifty years is as five eras, each shorter than the last, each ending when an open-source rival catches a closed-source pioneer.

The first era — recombinant DNA, roughly 1976 to 1985 — was the protein era. The scarce thing was the ability to make a human protein in a vat. Genentech (1976), Amgen (1980), and the cohort around them owned that ability, and they owned it for the better part of a decade because the equivalent of "publishing the model" did not exist — you could not upload a fermentation line to a preprint server. The craft lived in people and steel. The monopoly was long because the substrate was physical.

The second era was the bust — the mid-1980s through mid-1990s, when more capital had been poured into recombinant DNA than near-term products could absorb, the technique matured into general industrial use, the operators moved between companies carrying the craft with them, and the long tail of companies died. Genentech and Amgen survived because they had product cashflow. Regeneron was founded in 1988 into the wreckage. The lesson of the bust is the lesson of every cohort since: the technique stops being a secret roughly as fast as the people who hold it can be hired away.

The third era — roughly 1995 to 2010 — rotated the scarce thing from the protein to the platform for finding the next protein: antibody-discovery systems, high-throughput screens, structural-biology software. Roche bought Genentech in 2009 for $46.8 billion. Adimab was founded in 2007 to sell antibody discovery as a service rather than to build one company's pipeline. And the compression became measurable: ChemDraw, the closed cheminformatics suite, launched in 1985; RDKit, its open-source equivalent, shipped in 2006. Twenty-one years. That twenty-one-year gap is the number every subsequent era halves.

The fourth era — roughly 2010 to 2020 — turned the contract research organization from a back office into a discovery partner, and the gaps collapsed. AbCellera (2012) and Recursion (2013) were founded. Schrödinger's molecular-dynamics suite (1990) was matched by open tools around 2005 — fifteen years. Pipeline Pilot (2001) was matched by KNIME (2006) — five years. Benchling launched in 2012 and its open-source equivalents were essentially contemporary. In one generation the closed-to-open gap fell from twenty-one years to under one. The protagonist was not yet AI. It was the cloud and the open scientific stack — the infrastructure that let anyone, anywhere, assemble the substrate that used to take a regional cluster and years to build.

The fifth era is now, and AI is compressing what is left. AlphaFold 3 was matched in weeks. The 1980s recombinant-DNA cohort held a decade of monopoly because none of the mechanisms that close the gap operated at scale. The 2024 structural-prediction cohort held a quarter because all of them operate at full force at once: foundation models that are expensive to train and cheap to fine-tune, so the pretraining spills; cloud compute that flattens the infrastructure layer to something you rent by the minute; a talent layer that switches institutions every three to five years and carries the craft out the door; and composable open tools that assemble a closed suite's specific workflow in days. Those four forces compose, and the compression curve is what they produce when you run them together for forty years.

The era boundaries are fuzzy and the exact gaps depend on which products you pick. None of that matters to the shape. Across five eras, every fair comparison shows the lead time halving. Directionally certain, calendar-loose.

Figure 1 from When Biology Becomes Information

The price of compression: a hundredfold fall

The curve has a cost, and the cost is brutally legible if you plot the peak value of the platform leaders in each era. Each cohort captured roughly a tenth of the value of the cohort before it. Call it the 90 percent rule — not a law, but a regularity striking enough to deserve a name.

The first movers, the recombinant-DNA builders of the 1980s and 1990s, sit in the $100 billion band. Roche paid $46.8 billion for Genentech. Amgen is worth roughly $168 billion. Regeneron roughly $78 billion. The platform was a technique-plus-infrastructure stack — expression, purification, characterization, scaled fermentation, regulatory navigation — and the value captured at the platform level dwarfed any single molecule.

The service providers, the discovery-platform builders of the 2000s and 2010s, sit in the $10 billion band. Adimab's last private mark was around $10 billion. AbCellera went public in December 2020 at a peak intraday value of $15.66 billion. Twist Bioscience peaked above $9 billion. These are real businesses with real revenue and real customers, and as a cohort they cap an order of magnitude below the first movers.

The new entrants, the AI-discovery cohort of the 2020s, are pricing into the $1 billion band. None of the public pure-plays has crossed $5 billion; most cap below $2 billion. The private cohort has more upside still resident, but the public marks set the ceiling the market is currently willing to pay for the discovery layer in isolation.

Three eras, three orders of magnitude. And the single cleanest receipt that this is happening now, not in some theoretical future, is AbCellera. It went public on a covid-19 antibody it discovered and developed in roughly ninety days against an active outbreak, on the back of one of the strongest antibody-discovery platforms ever built. Peak value: $15.66 billion in December 2020. By June 2026: $1.97 billion. An 87 percent drawdown from the peak. The platform still works. The customers — Lilly, Regeneron, Pfizer, more than two hundred partners — still use it. What changed was not the technology. What changed was the market's willingness to pay a platform premium for a discovery engine. The premium did not disappear. It migrated.

The honest qualifier sits in the next sentence, because it always should. Public marks mislead. The AI cohort is young, its drugs are not yet approved, and the eventual value of a company that posts a Phase 3 readout in 2028 is not knowable from 2026 marks. But what you can say with confidence is that the platform-layer premium that built three $100 billion companies in the 1980s does not exist for the 2020s cohort at the same position in the stack. The question is no longer whether the premium migrated, but where it went — which is the subject of the back half of this piece.

Why the premium existed at all

Before tracking where the premium went, it is worth being precise about why anyone paid it, because the answer is the thing one who invest most needs to understand and the thing a founder most often gets wrong.

A platform is worth more than a single-asset drug company because it compounds R&D. The industry's own data, in research ZS Associates published in 2023, puts it at roughly seven subsequent assets per $100 million of R&D for validated platforms, against roughly four for non-platform companies. Two to one. The mechanism is shared infrastructure: the first program through a platform pays the fixed cost of building the expression line, validating the analytical methods, qualifying the manufacturers, training the regulatory team, writing the manufacturing-controls documentation. The second program inherits all of it. The fifth program costs roughly half what the first did. That is the flywheel, and the capitalized form of the flywheel is the platform premium.

Figure 2 from When Biology Becomes Information
Figure 3 from When Biology Becomes Information
Figure 4 from When Biology Becomes Information

The premium shows up in the business model as three revenue streams — an upfront payment when a partnership is signed, milestone payments when programs clear clinical and regulatory gates, and royalties when a partnered drug reaches the market. That three-part structure is not an accident of dealmaking; it is the platform diversifying across biotech's three valleys of death. The translational valley, where a candidate works in mice and fails in patients, is hedged by running a portfolio of partnerships at different stages. The financing valley, where burn outruns capital, is hedged by upfront cash. The regulatory valley, where a clinical hold freezes a pipeline, is hedged by milestone payments spread across programs. A single-asset company cannot diversify across any of these. A platform structurally can. It is running a venture portfolio on its own engine.

This matters in 2026 because the mix has rotated hard. Pharma deals are now roughly 70 percent contingent — the upfront fraction has shrunk, the milestone fraction has swelled. The platform that survives this shift is not the one that signs the most deals. It is the one whose engine is productive enough that the milestone density keeps the financing valley closed. The buyer has stopped paying a premium for access to a scarce technique, because the technique is no longer scarce, and has started paying for the ability to clear gates. The financing pattern rotated to match the commoditization of the chemistry layer. Founders who are still pitching the upfront are pitching the last era's deal.

Figure 5 from When Biology Becomes Information

The number of deals represented are all deals (acquisition, collaborations and licensing) with biotech platform companies. A company’s platform capability focus was defined as investment in platform biotech companies that allowed use and integration into large pharma R&D capabilities.

*ZS secondary research,2023 Whitepaper on therapeutic platforms

What is actually new in 2026

Three things are visible in the current data that the canonical 2023 framing could not, by construction, see. They are not independent — each one reprices the other two — but they are worth naming separately, because each one changes a different decision.

A fifth kind of company

The conventional way to classify biotech platforms is by defensibility — what the moat is made of. Modality platforms defend with composition-of-matter patents on the chemistry. Insight platforms defend with proprietary data or model weights. Services platforms defend with scale and switching costs. Vertical integrators defend with the cashflow from a commercial product. It is a good taxonomy, and it has organized a decade of thinking. But it has a blind spot, and the blind spot is exactly the company the present moment is producing.

Change the question. Stop asking what the moat is made of and start asking what the platform does. A fifth genus appears that the defensibility lens cannot surface: the closed-loop learning system, where every experiment retroactively improves every other target, where the primary balance-sheet asset is the dataset rather than the pipeline, and where the modality — small molecule, biologic, gene therapy — is an output the platform routes to rather than the thing the company is. Call it the Recursive Discovery Factory.

The defining inversion is this. In a modality platform, the molecule is the company — Moderna is mRNA, the chemistry constitutes the firm. In a Recursive Discovery Factory, the molecules in the clinic are evidence that the dataset is good. The asset is the loop. Recursion runs 2.2 million imaging experiments a week and embeds them into a single representation, so that an experiment on a fibrosis target also reprices every previous target in oncology, because the underlying representation does not know which disease you labeled it with. Roche's payments to Recursion are best understood not as funding a pipeline but as renting access to that representation. Bayer running Recursion's engine inside its own workflows is the same asset, externalized.

This is the only platform architecture that compounds as biology becomes information, because its moat is not a piece of biology that can be copied — it is the rate at which the system learns, which is harder to copy than any single model or molecule. That is the whole argument of this essay arriving at its point. When everything copyable gets copied in weeks, the only durable thing is the system that learns faster than it can be copied.

It is worth being honest about where this stands. On a maturity scorecard that a check-writer would actually use today, the Recursive Discovery Factory cohort loses to the vertical integrators, and loses on nearly every axis, because the vertical integrators have approved products and the factories mostly do not. The claim is not that the factories win today. The claim is that the economic shape — falling marginal cost of the next validated hit, dataset on the balance sheet — is the only shape that can punch through the $1–3 billion services ceiling without requiring a prior commercial product. Whether that shape resolves into something pharma has to pay up for is the empirical question, and it is dated: the next twenty-four months of clinical readouts decide it. Recursion's CDK7 readout, Insilico's idiopathic-pulmonary-fibrosis molecule moving to Phase 3 on the back of a positive Phase 2a in Nature Medicine — these are the selection step. If two of the clean cases post positive readouts, the architectural claim becomes a maturity claim and the cohort reprices upward. If none do, the loop is real on cost and false on translation, and the architecture stays useful while the valuation thesis weakens. That is the bet, and it is falsifiable on a calendar.

A buyer that is being forced to pay

The second change is not about the sellers at all. It is about the buyer, and it runs on a near-deterministic calendar that has nothing to do with sentiment.

Roughly $300 billion of branded pharmaceutical revenue loses patent exclusivity between 2025 and 2030 — about one-sixth of the industry's annual top line, concentrated in around 200 drugs and 70 blockbusters. Keytruda, Merck's $29-billion-a-year franchise and 42 percent of its revenue, loses US exclusivity in December 2028. Eliquis, worth roughly $19 billion combined to Bristol-Myers Squibb and Pfizer, erodes across 2026 to 2029. Novartis already lost Entresto, a $7.8 billion franchise, to generics in July 2025 — which is why Novartis led the 2025 acquirer table with more than $29 billion of deals.

This is the engine under the M&A wave. 2025 closed at roughly $228 billion in announced biotech and pharma transactions, with seventeen separate billion-dollar deals. And the pattern is legible: the buyer is almost always the pharma with the most visible cliff, the asset purchased is almost always late-stage and commercial-shaped, and pharma is now running its dealmaking like a venture portfolio — explicit allocations across modalities and stages, milestone-heavy terms that read like preferred-equity structures, post-deal operational support that looks like a fund's value-add. The buyer is no longer paying for the chemistry. The buyer is paying for cadence — the rate at which an engine throws clinic-ready candidates, because cadence is what fills a revenue hole on a deadline. The platform builder is being underwritten as a portfolio company by a pharma that runs a portfolio.

A geography that flattened

The third change extends the second one country at a time. The geographic monopoly that buttressed the value capture of the first three eras — the platform technique living in a dozen American and European institutions — is gone, because the substrate that produces drug candidates is now global. The same compute, the same open models, the same mobile talent layer sit in Shanghai, Hangzhou, Seoul, Bengaluru, and São Paulo.

China is the visible edge of this. Chinese-origin assets went from roughly 3 percent of Western pharma licensing in 2020 to roughly 30 percent in 2025 — a tenfold rise in five years, built on a substrate of post-2015 regulatory reform, lower trial costs, and a real domestic buyer market. The Western buyer pays the same dollar for an asset that reached clinical readiness in less time and at less capital. That is not a geopolitical surprise; it is a rational response to a structural shift.

India is the next question, and the answer is that India is not China, which is the most useful thing to understand about it. The biology talent exists, world-class at roughly a quarter of US wages. The regulator is moving — the March 2026 clinical-trial reforms narrowed the gap by an estimated ninety days per program. What is missing is the domestic buyer of scale: there is no Indian Hengrui, no Indian BeiGene, no $20 billion novel-mechanism acquirer. Which means the exit math is different — and possibly, counterintuitively, better. The services-platform ceiling that prices a US platform at $1–3 billion and reads as a death sentence at $200 million of Cambridge burn is a clean 30-to-60x return at $40 million of Bengaluru burn, if a buyer exists. The crack in the wall is Sun Pharma's $11.75 billion all-cash bid for Organon in April 2026 — the first time Indian pharma moved at scale to buy a US-listed innovative pipeline rather than be bought. One transaction does not make a buyer class. It makes a blueprint. Whether anyone follows is, again, a dated and falsifiable question.

What this means for you

The synthesis is only useful if it changes a decision. Here is the translation, by seat.

If you are building a platform, stop building your moat out of chemistry. The closed-to-open gap means whatever technical edge you have in your modality is measured in weeks to months, not years, and you will spend your whole company life watching the open-source stack catch you. The durable thing is not the molecule and not the model — both get copied. The durable thing is the loop: a closed system where every experiment improves the value of every other experiment, where the dataset is the asset and the molecule is the receipt that the asset works, and where your marginal cost of the next validated hit falls as the platform runs. If you cannot point to that falling-cost curve, you are not building a platform that survives the era. You are building a tool, and tools get priced as tools. Build the cadence, because cadence is what the buyer is now paying for.

If you are a VC, underwrite architecture, not modality. The question that organized the last decade — "what is the chemistry?" — is the wrong unit of analysis now, because the chemistry commoditizes. The right question is: what is the closed loop doing, on what substrate, and does the cost of the next validated hit fall as it runs? That reframing changes what you diligence, what you index founders on, and what you wait for. The selection step is public and dated — the clinical readouts of the next twenty-four months are the assay that tells you whether the Recursive Discovery Factory shape resolves into a maturity story or stays a promising architecture priced at the services ceiling. Do not pay the maturity multiple before the readout. Do not miss the readout when it comes.. The single most important fact is the one this whole piece has been documenting: the platform-layer premium has been repriced downward by a factor of roughly a hundred across two generations, and the third generation is pricing into another factor of ten. That does not mean biotech platforms are a bad allocation. It means the position in the stack where value accrues has moved, and a fund still underwriting 1980s platform multiples on 2020s discovery-layer businesses is paying for a premium that the market has already migrated away from. The durable value lives in cadence and the loop, and in the geographies where the cost base inverts the ceiling into a return. Ask your managers which question they are underwriting — chemistry or architecture — and whether they have a dated thesis for the readouts that resolve it. The ones who can answer that are reading the era correctly. The ones who cannot are buying the last one.

The honest close

Everything above is built on a foundation worth stating plainly, because a framework that cannot be wrong is not a framework. We will be wrong on the slope. The data is too noisy, the cohorts too small, the public marks too few to be quantitative about the exact multiples. We will probably be wrong about which Recursive Discovery Factory sub-type compounds fastest, about whether Indian platforms route through this architecture on the timeline the math suggests, and about at least one of the falsifiable predictions that sit underneath this worldview. The point of writing the bets down with dates attached is so the wrongness is legible and the next cut is sharper.

What I’m willing to commit to before the next year of data lands is the structure. The compression curve runs — the lead time a platform holds has fallen from a decade to a quarter, and it has fallen because biology keeps becoming information. The hundredfold fall is real and structural — each cohort captures a tenth of the last. The taxonomy is one continuous refinement, from output to defensibility to architecture, and the fifth genus was always going to appear once you asked what the platform does. The buyer has rotated to a venture posture and is paying for cadence. The geography has flattened, and the value-capture ceiling that depended on geographic monopoly is being repriced country by country.

And underneath all of it, the one idea. Biology is becoming an information technology. When something becomes information, it stops being scarce, and the value migrates to whatever is still scarce — which, in the end, is not any molecule, any model, or any technique, because all of those get copied. The only thing that stays scarce is a system that learns faster than it can be copied. Fifty years of the curve have been the slow proof of that sentence. The next twenty-four months are the test of where it leaves us.

The framework is the gift. The future is open. Let us find out together.