Drug discovery faces mounting pressure to deliver faster, cheaper results. But despite AI's promise, a critical bottleneck remains: the data itself. Paul Belcher, director of protein research strategy at global life sciences company Cytiva, explained in an interview with MIT Technology Review that AI models are hitting what he calls a data wall. 'The main cost in drug discovery is still the clinical phase,' Belcher said. 'AI is one approach that drug companies hope will not only save time but enable better quality candidates to reach the clinic.'

What You Need to Know

AI models in drug discovery rely on public datasets that lack negative results. Without access to failed experiments, models cannot learn to avoid bias. This places pressure on lab teams to validate more complex AI-generated compounds. Closing the data loop requires new data-sharing practices and integration with lab systems.

The Data Wall in AI-Driven Discovery

Since the 1950s, the cost of developing new drugs has roughly doubled every nine years, a trend known as Eroom's Law. Today, bringing a new drug to market takes an average of 10 to 15 years and costs between $1 billion and $2.5 billion. AI has become the industry's biggest bet on improving success rates and shortening timelines. Belcher noted that AI enables predictive design: companies can now use AI to design drug candidates from scratch and predict interactions before committing to research and development.

That shift, however, has created a new challenge. Traditional screening workflows were built to identify hits at scale using binary, low-fidelity techniques. AI generates a larger volume of more diverse candidates, demanding higher-throughput validation technologies. Belcher explained that AI 'increases the number of hits you get and potentially gives you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to validate and characterize those hits.'

  • Data wall: Public datasets are nearly exhausted, leading to diminishing returns for model improvement.
  • Publication bias: Most published results are positive, skewing model training toward success patterns.
  • Validation gap: AI-generated candidates still require physical testing, placing new strain on lab resources.

Why This Matters

The lack of negative data creates a fundamental problem for AI performance. Without access to a broad range of training data that includes failures, models cannot learn to avoid bias. Belcher noted that 'most publicly available datasets and scientific publications focus exclusively on positive results. This bias is almost like having one hand tied behind your back.' The result is an AI system that may replicate the same blind spots as traditional approaches, undermining its potential to truly accelerate drug discovery. Investors and regulators are watching closely: if AI cannot prove it reduces late-stage clinical failures, the billions poured into these tools may not pay off. Closing the data loop is not just a technical challenge. It is an economic and strategic imperative for an industry that can no longer afford 90 percent failure rates.