raminderpalsingh.com
caffeinated.bio mark: a cup of coffee whose steam forms a molecule

caffeinated.bioBiotech-in-a-Box

A private, always-on appliance for the computational steps of early drug discovery. It reads public data and literature, builds its own curated and computed layers, reasons over them with a local open-weight model, and maintains a ranked pipeline of program opportunities that updates as the evidence moves. One machine, which a research group owns outright.

It covers the computational steps. It measures nothing, and says so in every output it produces.

public dataPublic databases and literature only. No wet lab, no proprietary input.
one machineAn NVIDIA GB10 desktop system. 128 GB unified memory, one accelerator.
local modelA free open-weight model served on the appliance. No cloud model, no per-call fees.
always onContinuous ingest, change detection, and re-reasoning only where evidence moved.
nothing leavesResults, program state and reasoning stay on the machine. Only lookups go out.

The goal

A valid Biotech-in-a-Box approach to the computational steps in drug discovery, delivered as a private, easy-to-consume appliance.

The appliance is always on and always learning. It leverages targeted, effective practices in four areas: bioinformatics, cheminformatics, database layers, and local agentic engines. Each of those is a mature field on its own. What is new is running them together, continuously, on one private machine that a research group can own outright.

The intent is disruptive. If the computational half of early discovery runs on an appliance rather than a platform subscription or a cloud tenancy, the economics of investigating an early pipeline change: a group can run more opportunities in parallel, keep every piece of reasoning in-house, and scale investigation without scaling headcount or spend. That is the opportunity this project exists to demonstrate, and it is why the appliance form matters as much as the science inside it.

What it does not do is replace the laboratory. The appliance covers the computational steps. The criteria that decide a drug candidate are measured, not computed, so every proposal it produces ends in a costed list of the experiments that still stand between that program and a candidate.

What it is

One machine holding the data, the models and the reasoning needed to take a program from public evidence to a documented proposal.

What comes out

For each program, a proposal: a target and indication thesis with its evidence graded, a genetic direction-of-effect call, a competitive and patent position, a tractability read, chemical starting points with predicted properties and liabilities, a costed list of the experiments still needed, and a falsifier saying what would most move its rank. Across programs, a ranked pipeline with a change history.

The four practices

Each chosen for what it contributes, not for coverage.

BioinformaticsCausal genetic evidence rather than association scores alone: Mendelian randomization and colocalization over public GWAS, eQTL and pQTL data, giving a direction-of-effect call for each target and indication pair. Pocket and structure work reads existing PDB and AlphaFold structures rather than folding anything.
CheminformaticsCuration first: per-target structure-activity tables harmonized by assay type, with censored values handled and transcription errors flagged. Then property, ADMET and off-target models with an applicability-domain estimate on every prediction, generative design constrained by retrosynthesis so the output is makeable, and an exposure estimate that fills missing values from compound properties rather than stopping at a gap.
Database layersSix layers flowing one way: dated immutable mirrors, curated tables, computed tables marked as predicted, a knowledge graph, program state, ranking history. Every edge in the graph carries its provenance and its class, and is bitemporal: it records when it became true and when the appliance learned it.
Local agentic enginesfred does the evidence work: assembling what is known, tracing every claim to a source, flagging contradictions, checking its answers against what it retrieved, and writing the reasoning a person reads. It produces no number that enters the ranking. The ranking is arithmetic over stored, versioned inputs.

The division in the last two is the one that makes the output defensible. A reasoning model writes prose and can be argued with; it does not get to decide an order. Everything numeric comes from an engine with a version, an applicability domain and a stored input, so a ranking can be re-run from the record months later and checked.

Always on, always learning

Continuous, but not wasteful: only what changed is reasoned over again.

Ingestion runs on a schedule. Change detection works out what actually moved. Only the programs whose evidence changed are re-reasoned. The ranking recomputes cheaply because it is arithmetic; the written reasoning regenerates only where it has to. This matters on one machine, where reasoning throughput is the binding limit.

What triggers a re-reasoning

EventEffect
new publicationClaims extracted, traced and checked against the program's existing evidence; contradictions surfaced rather than averaged away.
trial changeA registration, a status change or a termination moves the competitive position and, for a failure, the prior on the mechanism.
new patent familyChanges the whitespace and sometimes the chemical starting points.
new structureA deposited structure can move a target from predicted-only to experimentally grounded.
safety signalA post-marketing signal on related chemistry re-weights the liability read.
source releaseA new release of a public database re-runs the curated and computed layers that depend on it.

What "always learning" means here. The store grows and the reasoning over it is refreshed: new evidence, new computed estimates, new rankings, and a record of which past predictions held. It does not mean the model changes. The model is pinned on purpose, because a result that cannot be reproduced cannot be defended.

Project status

As at 7 October 2026.

scope set . design written . appliance in place . build not started

The scope is set and the design is written. The target machine is in place. fred is specified and developed on separate hardware; this project consumes it as a locked vendor service and never changes it. No engine, database layer or ingest has been built yet, and the first engine is named below rather than assumed.

DecisionWhat it commits to
public dataPublic databases and literature only. No wet lab, and no proprietary data on the way in.
one applianceOne machine, no cloud. Nothing computed on it leaves it.
no model numbersfred reasons and writes; it never produces a number that enters the ranking.
reproducible rankThe ranking is arithmetic over stored, versioned inputs, re-runnable from the record alone.
graph as queryThe knowledge graph is an integration and query layer, not a link-prediction model.
co-foldingProtein-ligand co-folding is conditional on a benchmark per target class, not a standing engine.
acceptance testA dated prediction, written before any experiment, is what the appliance is judged on.
Open questionWhat decides it
first indicationAn unmet need with a test system available, not whichever area the method finds convenient. A per-area backtest can inform it but does not settle it.
patent chemistryWhether structures extracted from patents are in scope at the start, or whether patents enter as bibliography and claims text only.
exposure engineThe credible open PBPK option is a GPLv2 .NET suite whose Arm support is unverified. A property-based ADME estimator may replace it.
graph storeThe most natural embedded graph engine was archived in late 2025. A maintained fork, a different engine, or Parquet plus an index.

The identity and normalization engine, with bioactivity curation.

Three reasons it comes first. Every other engine, the knowledge graph and the backtest rest on resolving proteins, genes, compounds, salts, diseases and assays to stable identifiers, and that is where errors enter. Its quality can be measured against public data without making any modeling claim. And both it and the curation that follows are CPU-bound, so neither competes with the model for the one accelerator.

The contrary view, stated because it is a real cost: neither produces anything demonstrable on its own. An early visible result, such as a causal-genetics ranking for one indication, would show more sooner and rest on weaker foundations.

ItemState
scopeSet: the computational steps of early discovery, public data, one appliance.
designWritten, including the layer model, the engine set and the acceptance test.
applianceIn place. Platform characteristics measured from published specifications, not yet on the machine.
reasoning modulefred, developed separately, consumed here as a locked vendor service.
enginesNone built.
database layersNone built. No mirror has been taken.
backtestSpecified, not run.

How it fits together

Four columns: public sources, the six database layers, the work engines, and the output.

Engines both read the database layers and write back to them, which the double-headed arrows mark. The order of the engines is a reading order, not a fixed sequence: change detection decides which engines run for which program, and a scheduler decides when, because they share one accelerator.

The appliance: sources, database layers, engines and outputOne appliance: public sources in, a ranked pipeline outTHE APPLIANCE: one machine, 128 GB unified, one accelerator. Nothing leaves it.PUBLIC SOURCESDATABASE LAYERSWORK ENGINESOUTPUTGenetics and targetsOpen Targets, GWAS Catalog,GTEx, eQTL / pQTL, ClinVar,gnomAD, DepMap, HPALiteraturePubMed, Europe PMC openaccess subset, bioRxivChemistry, bioactivityChEMBL, PubChem, BindingDB,OpenADMET, ASAP andFragalysis, Enamine REALStructures, pathwaysPDB, AlphaFold DB,Reactome, WikiPathwaysClinical and safetyClinicalTrials.gov, EU CTR,openFDA, FAERS, Tox21PatentsUSPTO bulk, EPO OPS,Google Patents, SureChEMBL1 Raw mirrorsDated snapshots,tiered to fit the disk2 Curated tablesNormalized ids,harmonized assays3 Computed tablesEngine output, markedpredicted, with itsversion and range4 Knowledge graphEvery edge carriesprovenance and class;bitemporal; weighted5 Program stateContext files for fred:Markdown and JSONwith a README6 Ranking historyEvery ranking withits inputsOrchestrationChange detection: re-reason only wherethe evidence movedScheduler: one accelerator, jobs queueIdentity and normalizationResolves proteins, genes, compounds,diseases and assaysBioinformaticsCausal genetics: MR, colocalization,direction of effect (CPU only)Pocket and structure: reads PDB andAlphaFold; fpocket. Folds nothingCo-folding: conditional on a benchmarkCheminformaticsBioactivity curation: SAR tables (RDKit)ADMET and off-target: Chemprop modelsDesign: REINVENT4, AiZynthFinderDocking fallback: Vina or gninaExposure: ADME and PBPK estimatefred, on a local open-weight modelEvidence assembly, claims with provenance,contradictions, grounding checks, rationaleWrites no ranking numberRankingArithmetic on orthogonal axes;weights set by a personBacktest, then prospective testRetrodicts a frozen past corpus, thenpredicts before the assay is runPer programTarget and indication thesis,graded by evidenceGenetic direction of effectCompetition and patentsTractability, structural readChemical starting pointswith predicted propertiesCosted list of experimentsstill neededFalsifier: what wouldmove its rankProvenance on every claimPortfolioRanked pipeline with historyPrediction recordAlerts when evidence movesNot a drug candidateComputed estimates, notmeasurements. The gap listsays what must be measured.
Figure 1. The appliance. Public sources on the left feed six database layers; the work engines read and write those layers; the output is a proposal per program and a ranked portfolio. The dashed outline is the machine: everything inside it stays on it.

The appliance

An NVIDIA GB10 desktop system: 128 GB of unified memory, a single accelerator, 1 to 4 TB of storage, on an Arm platform.

What one appliance buys

What one appliance costs

Throughput figures are calculated from published memory bandwidth and corroborated by a third-party load test on the same platform. Nothing has been measured on the machine itself.

Proving it

A reproducible ranking makes the order auditable. It does not make it right.

The hard objection to a system like this is not whether it can assemble evidence. It is the long tail it produces, where many theses look similar and none is clearly better than the rest. The answer is falsification.

A retrospective backtest, freezing the corpus at a past date and checking whether the ranking retrodicts what later advanced or failed, is necessary, and it is cheap here because every edge records when the appliance learned it. But a backtest predicts yesterday's weather. The defining test is prospective: the appliance writes down its ranked predictions, with a date, before any experiment is run, and a laboratory test finds out whether they hold. One compound on one cell line is enough to make the claim real.

From public evidence to a prediction that can be proved wrongHow a claim from public data becomes something that can be proved wrong1. Evidence inPublic records only. Everyitem keeps its source andthe date the box learned it.2. Computed estimatesEngines add what nobodyhas published. Each outputcarries its engine versionand whether the inputwas in range.3. Reproducible rankingArithmetic over storedinputs. fred writes thereasoning but no number.Re-runnable from therecord alone.4. Dated predictionWritten down, with a date,before any experimentruns. Each one names theresult that wouldrefute it.5. ExperimentRun in a laboratory. Onecompound on one cell lineis enough. The appliancedid not and cannot runthis step.The result goes back into the store, right or wrongA refuted prediction is kept with the measurement that refuted it, and re-weights the ranking.Negative results are data: inactive compounds are stored, not discarded.A backtest alone predicts yesterday's weather. Stage 4 before stage 5 is what makes the claim falsifiable.
Figure 2. The appliance covers stages one to four. Stage five is a laboratory test, which is the point at which the output stops being an opinion. Results return to the store whether they confirm or refute.

This is also why negative results are kept. Falsifying a prediction needs negatives, and public negative data is thin, so inactive results are stored as data rather than discarded.

The choices, and what each is worth

Every line is a choice the appliance form forces or invites. The right-hand column is the point.

ChoiceWhat it costsWhat it is worth
public dataNo proprietary structure-activity data, and no measured results of any kind.Caps the output at a proposal with a gap list, and raises its credibility, because every claim traces to a public source.
one applianceOne accelerator for every job; no burst capacity.Pushes the design toward depth on few programs rather than breadth over many, and makes the result transferable as a copy.
indication focusThe portfolio is small and says nothing about the rest of the field.Raises the value of each proposal sharply. The stronger trade for one machine.
no model numbersMore engineering: every numeric input has to come from somewhere auditable.What makes the ranking defensible to a chemist. Without it the output degrades to a well-written opinion.
ligand-based firstLow-data targets get a weaker potency read.Prevents a predicted affinity that reads like a measurement from entering a proposal.
graph as queryNo headline claim that the system predicts new targets.Lowers the marketing value, raises the scientific value: graph link prediction is dominated by how well connected an entity already is.
dated predictionDelays the first defensible output and consumes machine hours.The difference between a ranked list and an evidenced one.

Built from open source

Open components where they exist, written from scratch where they do not.

The chemistry and data layers are well served on this platform. RDKit, PyTorch, OpenMM with CUDA, DuckDB, PyArrow, scikit-learn, LightGBM, XGBoost, gemmi and Biopython all publish Arm builds; Chemprop, AiZynthFinder and the structure-prediction tools are pure Python on top of them. REINVENT4 handles generative design, AiZynthFinder the retrosynthesis that keeps designs makeable, fpocket the pocket detection, and AutoDock Vina or gnina the docking.

Two components carry real risk, and both are named rather than hidden: physiologically based pharmacokinetic modeling, where the credible open option is a GPLv2 .NET-based suite whose Arm support is unverified; and the graph store, where the most natural embedded engine was archived in late 2025 after its company was acquired. Both are kept replaceable rather than central.

Licenses and Arm availability were checked from published package metadata. Nothing has been installed, built or run.

Where it works, and where it does not

The selector is the problem, not the method.

An indication with a real need for new treatment, where a prediction can also be tested. Within that, the appliance does better where human genetic evidence is strong, where the target class has enough public structures and ligand data to work with, where open experimental data is rich, and where competition is middling: enough prior work to supply data, not so much that the whitespace is gone.

It should do badly on CNS indications, where brain penetration and translation are poorly predicted; on Gram-negative antibacterials, where permeability and efflux are poorly predicted; on polygenic diseases with weak genetics; and on anything needing biologics. The standing tension is that where public data is richest, competition is usually highest.

An easy problem whose solution nobody needs is not worth solving. Choosing the indication by what the method finds convenient is the way this goes wrong, and it is a trap the field has fallen into before.

Known limits

In plain words.

It covers the computational steps only

It measures nothing and cannot select a drug candidate. Every proposal ends in a list of the experiments that still stand between it and one.

Public data is public to everyone

Anything findable in public data is findable by anyone holding the same data. The appliance works the unglamorous middle of the distribution systematically and shows its work. It does not find secrets.

Predicted is not measured

Every computed value is marked as computed, with the engine version that produced it and whether the input was inside that engine's applicability domain.

Private is not air-gapped

The appliance reads public databases over the internet, so the search terms it derives leave the machine. Its results, its program state and its reasoning do not.

Coverage has holes, and they are in the awkward places

The mirrors are tiered to fit the disk, so the long tail depends on public APIs being up. Those gaps fall in exactly the under-studied corners the appliance is meant to work, and they are reported in the output rather than hidden.

One machine

No burst capacity and no redundancy. A hardware failure stops everything.