Robust AI for Clinical Programs
This post is a short communication about a new publication I am starting that focuses on agentic analysis of clinical stage biotech and biopharma programs. If you are interested in this topic you can you check it out here:
The initial batch includes three posts from large, medium, and small cap companies.
Further reports will be published initially on a roughly weekly basis.
Brief Description
This Substack has recorded a history of developments at the intersection of AI and biology - some from the industry, some from the literature, and some of my own efforts. EmberFinch Research is a distillation of many of those efforts focused on the technically demanding field of clinical drug development assessment.
Life science research is often a somewhat provincial exercise — sitting at the intersection of academia and industry, the field is often split along the lines of science and business. This is reflected in the common cycling of early research teams out of companies that land on a development candidate and the accompanying funding reallocation.
However, from a macro industry perspective, these are inseparable. The ultimate success of any medicine is measured not solely of whether it has good science and clinical impact, but whether it has the business backing to be executed and financed to reach its potential. This is a non-trivial challenge for companies to navigate and indeed drug development investments are among the only types that precisely bake in failures on a wide range of variables — scientific risk, development risk, competitive risk, manufacturing risk, commercialization risk, and so on.
EmberFinch Research is designed to cover this scope in a concise manner.
Robust AI
LLMs are stochastic — they produce outputs based on trained statistical patterns and are known to frequently be confidently and convincingly wrong. This is a challenge when aiming to produce a concentrated, factually dense, and technically accurate output. But it is such outputs that are both desired and necessary in a business or enterprise setting.
The process is non-trivial. Most people have probably experienced using an LLM that provides citations only to find that the citations do not actually contain the information they are referencing. This is not just a hallucination issue, it is a fundamental issue of how the underlying computation is done. This is compounded by the need to encode complex or nuanced business logic or domain expertise.
However, despite these technical challenges, the implementation of robust AI systems will be an essential aspect of the future of business, scientific research, and many other use cases. While the underlying computation of LLMs as they exist today may not be able to eliminate all traces of hallucinations or stochastic responses, we have seen already that the architecture surrounding them can mitigate this to a very high degree of reliability.
Don’t Bet Against the Models
There is a common refrain, at least in the SF AI world, that advises against betting against the progress of models as underlying intelligences. I think this is a fair assessment, but I am not convinced that the inherent capabilities of base LLMs in their current incarnation will fundamentally improve in increasingly dramatic step functions. I am, however, convinced that the surrounding elements, better MoE architectures, more high quality data, robust harnesses, and improved tooling will continue a monotonic march toward increasing capability and reliability.
AI systems will be approximate co-workers. They will be very smart and capable, but just like humans, they will bring their slice of context to an interaction, and humans will bring a different context. For the foreseeable future, the context that humans bring cannot be replicated fully by AI — and I don’t expect this will ever be the case, just as the context of two humans will never fully overlap. The main question is whether humans start to substitute their knowledge — which will lead to overall degradation in critical thinking skills. Even with the most sophisticated coding agents, I find their proposed solutions are often quite poor and spending a moment to critically review the actual situations results in a solution that is both simpler and more effective than anything an AI might propose. Relying on autonomous coding agents to write any system of meaningful complexity is not a good idea, but they are great assistants.
The required robustness of an AI system will be dependent on the domain and consequences of the information it processes. It is fair to say that humans are not infallible either, and maintaining a bar of perfection is not practical or reasonable for AI systems. It is a believable argument against the replacement of humans by AI system, that a core function of employees is to provide a buffer of liability for decision makers and executives. Replacing employees with AI ultimately would require shifting the liability for errors upwards, and that may be unpalatable.
If you are interested in this topic or in an expanded library of clinical program reports, write to contact@emberfinch.xyz


