On-prem flavor prediction from chemical structure β taste and aroma, physicochemical behavior, formulation notes, and safety flags for any molecule, plus substitution search. Enter a common or IUPAC name (or a SMILES) and get a single, honest flavor read, running entirely on hardware you own.
Every head tells you how good it is. Each of the 195 trained heads publishes the decision
threshold it fires at and its measured out-of-fold precision β because AUROC alone will flatter a
model badly on rare classes. The ginger head scores AUROC 0.979 and is right 1 time in 10
when it fires; both are true, and only one of them was ever being reported. Heads that cannot be
right more than half the time ship marked indicative rather than confident: they keep their
score, their chips and every molecule they find, but they are never dressed up as answers.
How accurate is it, really? explains all of it without assuming an ML
background.
A single, confidence-tagged flavor read β taste, aroma, behavior, safety, and drop-in substitutions.
The interactive flavor-space map in 3D on real MW Γ logP Γ TPSA axes, colored by taste & aroma β every one of the 172 aroma + 6 taste classes labelled. 195 trained heads in all: 6 taste + 172 aroma + 5 mouthfeel + 12 safety.
8,869 unique molecules across the open datasets Β· taste + aroma + mouthfeel prediction from structure Β· a flavor library (start from a flavor β its character-impact molecule) and flavor designer (pick your notes β best food-safe molecules + drop-in swaps) Β· an interactive 2D/3D flavor-space map Β· 2D & 3D structure views. All on commercial-clean public data.
Per set (unique molecules): taste training 3,845 Β· aroma training 2,394 Β· odor corpus 2,255 Β· documented taste 676 Β· mouthfeel training 2,534 Β· Tox21 (safety) 7,823 Β· GRAS reference 2,781 Β· sweetness intensity 316 Β· character-impact aroma supplement 602 associations (open-gov-sourced). Every one of the 8,861 is enriched with names + measured properties from public-domain PubChem.
How the universe grows: the flavor-space map and enrichment table show 8,861 unique structures (deduped by connectivity skeleton), expanded by ingesting the full EU/GB flavourings Union List (~2,200 authorised, Open Government Licence v3) so the browse-able universe is food-forward. Meaningfully-distinct stereoisomers (e.g. R- vs S-limonene) that carry their own documented odor/taste are surfaced as first-class rows; nothing is dropped.
This is the commercial edition β Apache-2.0 and commercial-clean: every model (taste and aroma) trains only on permissively-licensed / public-domain data, so it's free to use, sell, and run behind your firewall. The commercial-clean aroma model ships here as presence/absence; a higher-fidelity academic edition adds the scored-intensity model from research odor data (NonCommercial terms) β see Two editions.
Status: the Python training pipeline and prediction core are built and tested; the .NET serving layer, React workbench, and packaging are in progress. See the roadmap β M0βM1 landed, M2βM6 next.
Given one molecule, Flavormancer returns a structured profile where every value is tagged by how it was derived, so nothing reads as more certain than its source:
- Taste β six trained heads: sweet / bitter / umami / sour / salty / tasteless (RandomForests on fingerprint + physicochemical features), plus a sweetness-intensity regressor. Sour and salty also keep a transparent chemistry rule (acid group / alkali-salt) as a deterministic cross-check alongside the model.
- Aroma β 172 odor-descriptor heads (citrus, floral, minty, almond, fatty, petroleum, earthy, medicinal, sulfurous, camphor, fruity, fishy, garlic, ethereal, ammoniacal, pungent, pine, rose, rancid, alcoholic, woody, green, grassy, putrid) trained on public-domain HSDB odor text + curated character-impact facts, surfaced for any molecule, each with its held-out CV-AUROC (0.71β0.98). Presence/absence, honestly β intensity is the "comes with your data" upgrade (no public intensity data exists). Documented odor + detection thresholds are shown where cited.
- Behavior β logP, molecular weight, TPSA, H-bonding, ring/atom counts (computed); water solubility (ESOL estimate); volatility tier and pKa ranges (qualitative); measured boiling point / vapor pressure when a property table is loaded (lookup β structure-based BP was evaluated and declined as too inaccurate).
- Mouthfeel β five trained chemesthesis heads (cooling / pungent / warming / astringent / tingling): the trigeminal sensation, trained on curated public-domain agents (menthol & WS-coolants, capsaicinoids, tannins, Sichuan-pepper sanshools), each with its held-out CV-AUROC. Distinct from the same-named aroma notes β the sensation, not the smell (menthol feels cool; WS-23 cools with almost no odour).
- Stability β oxidation / hydrolysis / photo watch-flags.
- Safety (defensive, caution-only) β a disclaimer + scope on every result, twelve Tox21 in-vitro assay heads (nuclear-receptor + stress-response, each with its CV-AUROC) surfaced as indicative review flags, structural tox-alert screening, a preliminary TTC/Cramer concern tier, an optional GRAS cross-reference, and EU declarable-allergen labeling. It flags for review; it never clears a compound for use.
- Formulation β a documented dangerous-mixture screen (benzene, nitrosamine, acrylamide, ethyl carbamate, furan, and more) and an OAV dosing-balance analysis that flags the component about to overpower a blend (quantitative when threshold tables are loaded).
- Substitutes & structural neighbors β two nearest-neighbor searches over the whole molecule universe for reformulation and cost-down: substitutes rank by taste + aroma profile match (a molecule that tastes and smells like the target β e.g. ethyl vanillin for vanillin β regardless of structure), while structural neighbors rank by Tanimoto/Morgan structure similarity (the look-alikes).
- Flavor Studio β one hub to pick any mix of everyday flavors (banana, saffron, pumpkin, bubble gumβ¦) and notes (citrus, floralβ¦) β ranked food-safe molecules + drop-in swaps. A flavor is a set of notes, so they live in one picker.
- Flavor-space map β every molecule embedded 2D/3D: a similarity layout (UMAP over fingerprints) and an interpretable property-axes layout (MW Γ logP Γ TPSA), colored by taste, aroma, or both.
- Chirality explorer β enumerates every stereoisomer (R/S and E/Z), surfacing the ones with distinct documented odor/taste (R- vs S-carvone) as their own reads.
- Master enrichment table β sortable/searchable grid of the whole universe with a "why it matters" note on every chemistry column.
- External references β per-molecule deep links to PubChem (public domain) and NIST WebBook for spectra (IR/MS/NMR) and GC retention indices, with an availability flag showing what each source has. We link, never rehost.
On aroma intensity (the one gated piece). The aroma heads above are presence/absence β which notes apply, from public-domain data. Scored intensity (how strong) is the marquee upgrade that comes with a customer's own panel data or a commercially-licensed set (Leffingwell PMP 2001); no public-domain intensity data exists. That's the only aroma capability still gated.
What aroma training needs from you: molecules (SMILES, or GC-MS to identify the compounds in your products) paired with your panel's expert odor descriptors (e.g. green / fruity / woody, ideally with intensity). GC-MS identifies the molecules; the sensory labels are what the model learns β GC-MS alone isn't enough.
For exactly what data unlocks each further capability (aroma, quantitative dosing, retention index) and the formats we accept from a client, see docs/DATA-REQUIREMENTS.md.
Scope: Flavormancer predicts flavor properties only. It is not a safety, toxicity, GRAS, regulatory, or stability determination. A prediction is never a clearance to consume.
Every output is labeled by how it was produced: computed (exact from structure), trained (ML on open data), rule (deterministic structural rule), estimate (published QSPR with known error), lookup (from a loaded reference table), and qualitative (a class/flag, not a number). Full map in docs/CAPABILITIES.md.
How does it all actually work? A plain-English tour of the method behind every feature β fingerprints, the taste/aroma models, the map, chirality, the mixture reactions β is in docs/HOW-IT-WORKS.md.
The value isn't any single number β those you can look up. It's four design choices:
- Prediction, not lookup. The trained models read a molecule from its structure, so they answer for a novel or unmeasured compound that's in no database β not just for known ones.
- On-premise, behind your firewall. Proprietary candidate structures never leave your network. You can screen confidential molecules without exposing your direction to any external service β something no public web tool can offer.
- One integrated read. Taste, behaviour, safety, and substitution in a single confidence-tagged screen, instead of stitching together half a dozen databases and manual checks per molecule.
- It sharpens on your own data. The open-data models are the floor. Trained on a user's own formulation and sensory data β on the same on-prem box β they become specific to that user's products, which no public dataset can be.
Flavormancer ships as two editions of one method:
| Flavormancer (this repo) | Flavormancer Research | |
|---|---|---|
| Edition | Commercial | Academic / open-source (coming soon) |
| License | Apache-2.0 | open-source, research / NonCommercial |
| Data | commercial-clean open data only | adds research odor datasets with NonCommercial terms |
| Aroma | 172 presence/absence descriptor heads ship (public-domain HSDB); scored intensity is trained on your data or a licensed set (PMP 2001) | full open model incl. intensity (research odor data) |
| Use | free to use, sell, run on-prem | research, teaching, advancing the method |
The split is deliberate. The richest aroma data is licensed for research only, so it can't ship in a product you sell β keeping it out is exactly what makes this edition clean to use and sell, and the academic edition is where that fuller model lives.
Python trains the models offline; a .NET application serves them at runtime β nothing at runtime depends on Python.
data sources ββΊ Python training (build-time) ββΊ ONNX (taste) βββ
RDKit Β· scikit-learn β
βΌ
React workbench ββ JSON API ββ ASP.NET Core + ONNX Runtime + Postgres/pgvector
β
βΌ
Docker Compose on a single on-prem box
| Layer | Technology |
|---|---|
| Model training (build-time) | Python Β· RDKit Β· scikit-learn Β· skl2onnx |
| Model handoff | ONNX |
| App / API | ASP.NET Core (C#) |
| ML serving | ONNX Runtime, in-process in .NET |
| Frontend | React |
| Database | PostgreSQL + pgvector |
| Deploy | Linux + Docker Compose (single box) |
training/ Python β dataset build + model training (build-time)
api/ ASP.NET Core β app, auth, endpoints, ONNX serving
frontend/ React β the workbench UI
infra/ Dockerfiles, docker-compose.yml, deploy
docs/ architecture, capabilities, and design docs
tests/ pytest suite for the prediction core
With Docker β the app plus a pgvector-backed Postgres, one command:
cp .env.example .env # every value has a working default
docker compose up -d
curl localhost:8000/healthzTrained models are not in the image (they are ~1 GB and change on every retrain) β point
MODELS_DIR at them and they mount read-only at run time. Set FLAVORMANCER_HOME if you run
without Compose.
From source β training/SETUP.md covers the install;
docs/DATA-PIPELINE.md is the clean-machine walkthrough with every
build step in dependency order, timings, and an end-to-end check that verifies a prediction
rather than just that the server started.
Datasets and trained models are not committed β the training scripts pull their sources and
.gitignore keeps artifacts out of the repo.
The Docker path has not yet been run end to end on a machine with Docker installed β the Compose file parses and the path handling is tested, but treat the first
docker compose upas the test rather than a guarantee.
Built by Echelon Technology Solutions β a small team where everyone is cross-training toward full-stack.
| Role | |
|---|---|
| Austin Β· @rvnminers-A-and-N | Founder & lead β training pipeline, backend, architecture, models, and roadmap; mentors the team across their areas |
| Jamie | Front-end lead β owns the UI/UX |
| Aaron Β· @Jabbacado | Full-stack (in training) β backend, infrastructure, AI, UI/UX |
| Ty Β· @Frostfire0101 | Full-stack (in training) β backend, infrastructure, networking, AI, UI/UX |
Austin β B.S. Chemistry (ACS-certified) & B.S. Physics, magna cum laude, Western Kentucky University (minors in Mathematics & Astronomy), plus graduate coursework in a computational-chemistry M.S. program. Background in molecular-dynamics simulation β protein conformational states and SOβ at the airβwater interface β on the WKU HPC cluster: the structure-to-property grounding behind Flavormancer.
See CONTRIBUTING.md for branching, commit conventions, and the review workflow. Commits are signed and signed off (DCO + CLA); PRs are small, linked to an issue, and squash-merged after review.
Licensed under the Apache License 2.0 β code license only; datasets and any pretrained models carry their own licenses, noted where used. The academic edition is a separate repository under its own (NonCommercial) terms.


