A private, always-on appliance for the computational steps of early drug discovery. It reads public data and literature, builds its own curated and computed layers, reasons over them with a local open-weight model, and maintains a ranked pipeline of program opportunities that updates as the evidence moves. One machine, which a research group owns outright.
It covers the computational steps. It measures nothing, and says so in every output it produces.
A valid Biotech-in-a-Box approach to the computational steps in drug discovery, delivered as a private, easy-to-consume appliance.
The appliance is always on and always learning. It leverages targeted, effective practices in four areas: bioinformatics, cheminformatics, database layers, and local agentic engines. Each of those is a mature field on its own. What is new is running them together, continuously, on one private machine that a research group can own outright.
The intent is disruptive. If the computational half of early discovery runs on an appliance rather than a platform subscription or a cloud tenancy, the economics of investigating an early pipeline change: a group can run more opportunities in parallel, keep every piece of reasoning in-house, and scale investigation without scaling headcount or spend. That is the opportunity this project exists to demonstrate, and it is why the appliance form matters as much as the science inside it.
What it does not do is replace the laboratory. The appliance covers the computational steps. The criteria that decide a drug candidate are measured, not computed, so every proposal it produces ends in a costed list of the experiments that still stand between that program and a candidate.
One machine holding the data, the models and the reasoning needed to take a program from public evidence to a documented proposal.
For each program, a proposal: a target and indication thesis with its evidence graded, a genetic direction-of-effect call, a competitive and patent position, a tractability read, chemical starting points with predicted properties and liabilities, a costed list of the experiments still needed, and a falsifier saying what would most move its rank. Across programs, a ranked pipeline with a change history.
Each chosen for what it contributes, not for coverage.
The division in the last two is the one that makes the output defensible. A reasoning model writes prose and can be argued with; it does not get to decide an order. Everything numeric comes from an engine with a version, an applicability domain and a stored input, so a ranking can be re-run from the record months later and checked.
Continuous, but not wasteful: only what changed is reasoned over again.
Ingestion runs on a schedule. Change detection works out what actually moved. Only the programs whose evidence changed are re-reasoned. The ranking recomputes cheaply because it is arithmetic; the written reasoning regenerates only where it has to. This matters on one machine, where reasoning throughput is the binding limit.
| Event | Effect |
|---|---|
| new publication | Claims extracted, traced and checked against the program's existing evidence; contradictions surfaced rather than averaged away. |
| trial change | A registration, a status change or a termination moves the competitive position and, for a failure, the prior on the mechanism. |
| new patent family | Changes the whitespace and sometimes the chemical starting points. |
| new structure | A deposited structure can move a target from predicted-only to experimentally grounded. |
| safety signal | A post-marketing signal on related chemistry re-weights the liability read. |
| source release | A new release of a public database re-runs the curated and computed layers that depend on it. |
What "always learning" means here. The store grows and the reasoning over it is refreshed: new evidence, new computed estimates, new rankings, and a record of which past predictions held. It does not mean the model changes. The model is pinned on purpose, because a result that cannot be reproduced cannot be defended.
As at 7 October 2026.
scope set . design written . appliance in place . build not started
The scope is set and the design is written. The target machine is in place. fred is specified and developed on separate hardware; this project consumes it as a locked vendor service and never changes it. No engine, database layer or ingest has been built yet, and the first engine is named below rather than assumed.
| Decision | What it commits to |
|---|---|
| public data | Public databases and literature only. No wet lab, and no proprietary data on the way in. |
| one appliance | One machine, no cloud. Nothing computed on it leaves it. |
| no model numbers | fred reasons and writes; it never produces a number that enters the ranking. |
| reproducible rank | The ranking is arithmetic over stored, versioned inputs, re-runnable from the record alone. |
| graph as query | The knowledge graph is an integration and query layer, not a link-prediction model. |
| co-folding | Protein-ligand co-folding is conditional on a benchmark per target class, not a standing engine. |
| acceptance test | A dated prediction, written before any experiment, is what the appliance is judged on. |
| Open question | What decides it |
|---|---|
| first indication | An unmet need with a test system available, not whichever area the method finds convenient. A per-area backtest can inform it but does not settle it. |
| patent chemistry | Whether structures extracted from patents are in scope at the start, or whether patents enter as bibliography and claims text only. |
| exposure engine | The credible open PBPK option is a GPLv2 .NET suite whose Arm support is unverified. A property-based ADME estimator may replace it. |
| graph store | The most natural embedded graph engine was archived in late 2025. A maintained fork, a different engine, or Parquet plus an index. |
The identity and normalization engine, with bioactivity curation.
Three reasons it comes first. Every other engine, the knowledge graph and the backtest rest on resolving proteins, genes, compounds, salts, diseases and assays to stable identifiers, and that is where errors enter. Its quality can be measured against public data without making any modeling claim. And both it and the curation that follows are CPU-bound, so neither competes with the model for the one accelerator.
The contrary view, stated because it is a real cost: neither produces anything demonstrable on its own. An early visible result, such as a causal-genetics ranking for one indication, would show more sooner and rest on weaker foundations.
| Item | State |
|---|---|
| scope | Set: the computational steps of early discovery, public data, one appliance. |
| design | Written, including the layer model, the engine set and the acceptance test. |
| appliance | In place. Platform characteristics measured from published specifications, not yet on the machine. |
| reasoning module | fred, developed separately, consumed here as a locked vendor service. |
| engines | None built. |
| database layers | None built. No mirror has been taken. |
| backtest | Specified, not run. |
Four columns: public sources, the six database layers, the work engines, and the output.
Engines both read the database layers and write back to them, which the double-headed arrows mark. The order of the engines is a reading order, not a fixed sequence: change detection decides which engines run for which program, and a scheduler decides when, because they share one accelerator.
An NVIDIA GB10 desktop system: 128 GB of unified memory, a single accelerator, 1 to 4 TB of storage, on an Arm platform.
Throughput figures are calculated from published memory bandwidth and corroborated by a third-party load test on the same platform. Nothing has been measured on the machine itself.
A reproducible ranking makes the order auditable. It does not make it right.
The hard objection to a system like this is not whether it can assemble evidence. It is the long tail it produces, where many theses look similar and none is clearly better than the rest. The answer is falsification.
A retrospective backtest, freezing the corpus at a past date and checking whether the ranking retrodicts what later advanced or failed, is necessary, and it is cheap here because every edge records when the appliance learned it. But a backtest predicts yesterday's weather. The defining test is prospective: the appliance writes down its ranked predictions, with a date, before any experiment is run, and a laboratory test finds out whether they hold. One compound on one cell line is enough to make the claim real.
This is also why negative results are kept. Falsifying a prediction needs negatives, and public negative data is thin, so inactive results are stored as data rather than discarded.
Every line is a choice the appliance form forces or invites. The right-hand column is the point.
| Choice | What it costs | What it is worth |
|---|---|---|
| public data | No proprietary structure-activity data, and no measured results of any kind. | Caps the output at a proposal with a gap list, and raises its credibility, because every claim traces to a public source. |
| one appliance | One accelerator for every job; no burst capacity. | Pushes the design toward depth on few programs rather than breadth over many, and makes the result transferable as a copy. |
| indication focus | The portfolio is small and says nothing about the rest of the field. | Raises the value of each proposal sharply. The stronger trade for one machine. |
| no model numbers | More engineering: every numeric input has to come from somewhere auditable. | What makes the ranking defensible to a chemist. Without it the output degrades to a well-written opinion. |
| ligand-based first | Low-data targets get a weaker potency read. | Prevents a predicted affinity that reads like a measurement from entering a proposal. |
| graph as query | No headline claim that the system predicts new targets. | Lowers the marketing value, raises the scientific value: graph link prediction is dominated by how well connected an entity already is. |
| dated prediction | Delays the first defensible output and consumes machine hours. | The difference between a ranked list and an evidenced one. |
Open components where they exist, written from scratch where they do not.
The chemistry and data layers are well served on this platform. RDKit, PyTorch, OpenMM with CUDA, DuckDB, PyArrow, scikit-learn, LightGBM, XGBoost, gemmi and Biopython all publish Arm builds; Chemprop, AiZynthFinder and the structure-prediction tools are pure Python on top of them. REINVENT4 handles generative design, AiZynthFinder the retrosynthesis that keeps designs makeable, fpocket the pocket detection, and AutoDock Vina or gnina the docking.
Two components carry real risk, and both are named rather than hidden: physiologically based pharmacokinetic modeling, where the credible open option is a GPLv2 .NET-based suite whose Arm support is unverified; and the graph store, where the most natural embedded engine was archived in late 2025 after its company was acquired. Both are kept replaceable rather than central.
Licenses and Arm availability were checked from published package metadata. Nothing has been installed, built or run.
The selector is the problem, not the method.
An indication with a real need for new treatment, where a prediction can also be tested. Within that, the appliance does better where human genetic evidence is strong, where the target class has enough public structures and ligand data to work with, where open experimental data is rich, and where competition is middling: enough prior work to supply data, not so much that the whitespace is gone.
It should do badly on CNS indications, where brain penetration and translation are poorly predicted; on Gram-negative antibacterials, where permeability and efflux are poorly predicted; on polygenic diseases with weak genetics; and on anything needing biologics. The standing tension is that where public data is richest, competition is usually highest.
An easy problem whose solution nobody needs is not worth solving. Choosing the indication by what the method finds convenient is the way this goes wrong, and it is a trap the field has fallen into before.
In plain words.
It measures nothing and cannot select a drug candidate. Every proposal ends in a list of the experiments that still stand between it and one.
Anything findable in public data is findable by anyone holding the same data. The appliance works the unglamorous middle of the distribution systematically and shows its work. It does not find secrets.
Every computed value is marked as computed, with the engine version that produced it and whether the input was inside that engine's applicability domain.
The appliance reads public databases over the internet, so the search terms it derives leave the machine. Its results, its program state and its reasoning do not.
The mirrors are tiered to fit the disk, so the long tail depends on public APIs being up. Those gaps fall in exactly the under-studied corners the appliance is meant to work, and they are reported in the output rather than hidden.
No burst capacity and no redundancy. A hardware failure stops everything.