raminderpalsingh.com
fred project mascot

fredv0.3.0

A research service that answers a scientific question by looking things up rather than by recalling them, and returns a record of every lookup it made.

0.3.0The released version, deployed as a sealed container.
wire 1.2Carried on every response from the service.
1 at a timeOne question in flight; no queue, no cancel.
not deterministicThe same question can take a different route to a different answer.

What fred is

The loop, and where it sits.

fred is a research assistant service that answers a scientific question by looking things up rather than by recalling them. It works as a loop run by a language model: it reasons about what to check next (Thought), calls one public database (Action: UniProt, PubChem, PubMed, Open Targets, Reactome, or a web search), reads the result (Observation), and goes round again until it decides it can answer or its step budget is spent.

Applications such as Compound Insights and SIM send fred a question and receive an answer with a record of every lookup. fred is a component inside them, not a standalone tool.

Ten tools are registered and all ten are offered on every turn. One language model backend is configured: kimi-k2.6:cloud, reached through a local Ollama daemon. The step budget defaults to 30 turns and the caller can set it between 1 and 200.

AAAWgmp1bWIAAAAeanVtZGMycGEAEQAQgAAAqgA4m3EDYzJwYQAAABZcanVtYgAAAEdqdW1kYzJtYQARABCAAACqADibcQN1cm46YzJwYTpkNDBlYWY2ZC1lNjYzLTQ4MmEtOGM2MS00M2ZlZDZiZWY4YmYAAAADl2p1bWIAAAApanVtZGMyYXMAEQAQgAAAqgA4m3EDYzJwYS5hc3NlcnRpb25zAAAAALxqdW1iAAAARGp1bWRjYm9yABEAEIAAAKoAOJtxE2MycGEuaW5ncmVkaWVudC52MwAAAAAYYzJzaP0ixTtPoXXnozYRmSv0tV8AAABwY2JvcqNpZGM6Zm9ybWF0bWltYWdlL3N2Zyt4bWxqaW5zdGFuY2VJRHgseG1wOmlpZDpkY2JiNGZkYy1mOTVlLTQ5MzctOTJiYi0wNjA1YmFjZmU3MDRscmVsYXRpb25zaGlwaHBhcmVudE9mAAAB4mp1bWIAAABBanVtZGNib3IAEQAQgAAAqgA4m3ETYzJwYS5hY3Rpb25zLnYyAAAAABhjMnNojd43yHFGJ2MTnAXhCJiBxAAAAZljYm9yomdhY3Rpb25zgqJmYWN0aW9ua2MycGEub3BlbmVkanBhcmFtZXRlcnOha2luZ3JlZGllbnRzgaJjdXJseC1zZWxmI2p1bWJmPWMycGEuYXNzZXJ0aW9ucy9jMnBhLmluZ3JlZGllbnQudjNkaGFzaFgg0GwBqpelQqbw5iiQ/YJk0crWD33xbwxbS0EKAuKWVoWkZmFjdGlvbngdY29tLmFudGhyb3BpYy5jbGF1ZGUucHJvdmlkZWRqcGFyYW1ldGVyc6F4H2NvbS5hbnRocm9waWMub3JpZ2luLWNvbmZpZGVuY2VndW5rbm93bmtkZXNjcmlwdGlvbnhmQ2xhdWRlIHByb3ZpZGVkIHRoaXMgZmlsZSBhdCB0aGUgcmVxdWVzdCBvZiBhIHVzZXIgYW5kIG1heSBoYXZlIGNyZWF0ZWQgb3IgbW9kaWZpZWQgdGhlIGZpbGUgY29udGVudHMubXNvZnR3YXJlQWdlbnShZG5hbWVmQ2xhdWRlcmFsbEFjdGlvbnNJbmNsdWRlZPUAAADIanVtYgAAAEBqdW1kY2JvcgARABCAAACqADibcRNjMnBhLmhhc2guZGF0YQAAAAAYYzJzaMHBm7Jr1fK8HusmKW1rE2AAAACAY2JvcqVjYWxnZnNoYTI1NmNwYWRNAAAAAAAAAAAAAAAAAGRoYXNoWCDBerKRO0Jnpg0EeI6vLSUq99da40SujO8VC7pXV0K+FGRuYW1lbmp1bWJmIG1hbmlmZXN0amV4Y2x1c2lvbnOBomVzdGFydBjBZmxlbmd0aBkeBAAAAj5qdW1iAAAAJ2p1bWRjMmNsABEAEIAAAKoAOJtxA2MycGEuY2xhaW0udjIAAAACD2Nib3KlY2FsZ2ZzaGEyNTZpc2lnbmF0dXJleE1zZWxmI2p1bWJmPS9jMnBhL3VybjpjMnBhOmQ0MGVhZjZkLWU2NjMtNDgyYS04YzYxLTQzZmVkNmJlZjhiZi9jMnBhLnNpZ25hdHVyZWppbnN0YW5jZUlEeCx4bXA6aWlkOjY0MTA3MDI4LWE4ZjctNDUzYy05Mjk0LWJlMzEyMzEzNjE0N3JjcmVhdGVkX2Fzc2VydGlvbnODomN1cmx4LXNlbGYjanVtYmY9YzJwYS5hc3NlcnRpb25zL2MycGEuaW5ncmVkaWVudC52M2RoYXNoWCDQbAGql6VCpvDmKJD9gmTRytYPffFvDFtLQQoC4pZWhaJjdXJseCpzZWxmI2p1bWJmPWMycGEuYXNzZXJ0aW9ucy9jMnBhLmFjdGlvbnMudjJkaGFzaFggC9z9wXwP5b+qyvMHytKAdK4eZneS4eunoYcPJdXvTHyiY3VybHgpc2VsZiNqdW1iZj1jMnBhLmFzc2VydGlvbnMvYzJwYS5oYXNoLmRhdGFkaGFzaFggr60fWEQdN9XSQtC/Q2KOt+s+ZixVWTXsJPh18R4cDyx0Y2xhaW1fZ2VuZXJhdG9yX2luZm+jZG5hbWVvQW50aHJvcGljIEZpbGVzZ3ZlcnNpb25lMS4wLjBrc3BlY1ZlcnNpb25lMi40LjAAABA4anVtYgAAAChqdW1kYzJjcwARABCAAACqADibcQNjMnBhLnNpZ25hdHVyZQAAABAIY2JvctKEWQISogEmGCFZAgowggIGMIIBjaADAgECAhRA5aAK7sI50L64g/oGQgU9Z1UTADAKBggqhkjOPQQDAzBJMRcwFQYDVQQKEw5BbnRocm9waWMsIFBCQzEuMCwGA1UEAxMlQW50aHJvcGljIENvbnRlbnQgQ3JlZGVudGlhbHMgUm9vdCBDQTAeFw0yNjA4MDcxODQzNTZaFw0yODA4MDYxOTQzNTZaMEQxFzAVBgNVBAoTDkFudGhyb3BpYywgUEJDMSkwJwYDVQQDEyBBbnRocm9waWMgQ2xhdWRlIENvbnRlbnQgU2lnbmluZzBZMBMGByqGSM49AgEGCCqGSM49AwEHA0IABJh6CmvLUBgFFNU0vUKlOVtE6djd17L5SuwX0LemFisBM3dkd/3cyjxFA3Qo5S46fX0/ihY0VZ7mfb9KF703t5OjWDBWMA4GA1UdDwEB/wQEAwIHgDAVBgNVHSUEDjAMBgorBgEEAYPoXgIBMAwGA1UdEwEB/wQCMAAwHwYDVR0jBBgwFoAUzlHiBIFOZFsj+OPEz5o+nMHXXMIwCgYIKoZIzj0EAwMDZwAwZAIwMXMdFJ4BetLLVY7ORuE9noqbbAZOZn/aArXyTwFAZfKrPzxF2vPoJNf1+UCdg1XGAjBwX1zd9WGqYkqmL5SFqw1QySjr1zJfpJM9+1rdDwSPLMOPOjKuiXjoU/pUUeG9RwmhY3BhZFkNngAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAPZYQEsZGcKiyINUf1V0PKzPKCCIW4gBiJllsljTpRTwpdtteZW4IjCjwDDJw9jWVDuDYmv5eBOPp3HfLDkNQ0Bdqu0= Your application (Compound Insights, SIM) 1. asks a question 5. reads the final answer and its provenance record (or the step budget is spent) fred v0.3.0: a ReAct agent the model (kimi-k2.6 via Ollama) decides each step 2. THOUGHT what do I know? what should I check next? 3. ACTION choose one of ten tools and its arguments 4. OBSERVATION read the tool's reply every step recorded: tool, arguments, time, outcome, size and fingerprint of the reply Public databases UniProt (proteins) PubChem (compounds) PubMed (literature) Open Targets Reactome (pathways) Web search Golden test suite (separate harness, run against the live service) tools reachable > citations checked against database replies > independent AI judge > compared with a stored baseline question final answer query reply loop
Figure 1. The ReAct loop: Thought, Action, Observation, and back, with every step recorded; the final answer leaves the loop from Thought; the golden suite checks the live service.

What v0.3.0 adds: provenance

Seeing, after the fact, exactly what fred did to reach an answer.

What you get back

The status record, and the eleven fields kept for every tool call.

A completed run returns the model's final text as final_response, together with tool_calls: one record per tool call, in call order, complete. That list is the provenance record and the reason this release exists.

Eleven fields per tool call

FieldWhat it holds
toolThe name of the tool that was called.
argsThe parsed arguments. Malformed model JSON arrives as an empty object, which cannot be told apart from a call with no arguments.
result_summaryAn excerpt: at most 500 characters of the body, with an ellipsis when cut. Not the result.
truncatedTrue only when the body was cut before the model saw it, at a cap of 16,000 characters.
result_bytesUTF-8 length of the complete body as the tool returned it.
result_sha256SHA-256 of that complete body.
iterationZero-based index of the turn the call was made on.
started_atISO 8601 UTC timestamp.
duration_msWall clock for the call; -1 when unmeasured.
is_errorTrue when the tool returned an error outcome, raised, or was not registered; false on found and on found-nothing.
resolved_pathAbsolute path for the two file tools, set even when the workspace refused; null otherwise.

The complete body is not retained anywhere. Only its length and its hash survive. Recovering the full result of a lookup means asking the database the same question again.

The outcome vocabulary

Every tool body is one JSON object whose first key is outcome, with three possible values: OK, EMPTY (an authority answered "nothing"), and ERROR. An error body also carries an error_class, one of twelve: NO_RESPONSE, UPSTREAM_STATUS, UNEXPECTED_CONTENT_TYPE, UNPARSEABLE_BODY, UNEXPECTED_SHAPE, UPSTREAM_REPORTED_ERROR, INVALID_ARGUMENT, NOT_CONFIGURED, PATH_UNRESOLVABLE, WORKSPACE_REFUSED, TARGET_NOT_FOUND, TARGET_UNUSABLE.

is_error is two-valued and the vocabulary is three-valued: found-nothing is not an error, and the flag does not say found-nothing. Read outcome to tell "found nothing" from "found".

Key fields on the status record

FieldWhat it holds
wire_versionAlways "1.2".
run_id, tenant_idThe run's UUID, and the caller's identifier as sent. The tenant identifier is logged, not enforced.
statepending, running, completed, or failed. There is no cancelled state.
iterationsCount of completed turns. Live during the run, lagging one turn; exact on the terminal record.
tool_call_countCalls made so far. Live, lagging one turn. Counts calls, not successes, and survives a failure.
tools_offeredThe ten tool names in registry order, or null until the run is terminal.
prompt_sha256SHA-256 of the prompt as received. The prompt itself is never echoed back.
samplingRecords that fred set no sampling parameter: temperature, top_p and seed are all null.
modelProvider, model, model class, and the tokenizer name and revision. Null while the run is pending.
created_at, updated_at, elapsed_msISO 8601 UTC timestamps, and a monotonic elapsed time that never decreases across polls.
resultPresent on completed, null otherwise.
errorPresent on failed, null otherwise. Its kind is one of max_iterations_reached, llm_error, or internal. A tool failure does not fail a run.

A failed run carries its tool call count but no list of tool call records: the calls were made, and their records are not on the wire.

How it was tested

Three real questions, asked of the deployed service, checked four ways.

fred's own test set runs 936 tests, all passing. Every database tool has tests that a failing server, a timeout, or a malformed reply is reported as a failure. Those tests exercise the code; they say nothing about whether an answer is right. That is what the golden test suite is for.

The three questions

Each question was chosen so that a correct answer is a small set of specific values that a database holds today, and so that a wrong answer is unambiguous rather than a matter of opinion. Each is asked three times, because fred is not deterministic and one run proves nothing.

The expected answers were not written from memory. For each case the database was queried directly from the host, the reply was saved, and its size and SHA-256 recorded in the case file alongside the endpoint that produced it. The case file is fixed before the question is ever put to fred, and a case is never edited to match a result.

3 of 3 passed every check, 2026-09-09.

The question. What is the UniProt accession, the primary gene symbol, and the sequence length in amino acids of the human protein superoxide dismutase [Cu-Zn]? State the organism. Cite the UniProt accession in your answer.

Must be rightValue
accessionP00441
gene symbolSOD1
sequence length154 amino acids
organismHomo sapiens

Counted wrong. Any other length, any other accession for the human entry, any other primary gene symbol, or the claim that the human SOD1 entry is unreviewed. The accession must be cited, not merely mentioned.

Which tools. The UniProt protein lookup must fire. The web search and both file tools must not: the answer has to come from the database, not from the open web.

Ground truth. Measured 2026-09-07 against the UniProt REST search endpoint; a 1,411-byte reply saved with its SHA-256, giving accession P00441, gene SOD1, length 154, organism Homo sapiens, entry type reviewed (Swiss-Prot).

3 of 3 passed every check, 2026-09-09.

The question. A 2026 primary research article reports that TARDBP mediates a MAP3K11/SLC3A2/GPX4 axis in a rat model of Alzheimer's disease by enhancing the stability of KRAS mRNA. Find that article in PubMed. State its PubMed ID, its DOI, the journal, the publication year and the surname of its first author. Cite the PMID in your answer.

Must be rightValue
PMID42144687
DOI10.1111/jcmm.71181
journalJournal of Cellular and Molecular Medicine
year2026
first authorZhao

Counted wrong. Any other PMID, DOI, year, journal or first author for this article, and the claim that the study organism was human or mouse rather than rat. Declining to answer, or reporting that the article could not be found, is also wrong: the article exists.

Which tools. The PubMed literature search must fire; the web search and both file tools must not.

Ground truth. Measured 2026-09-07 from the PubMed E-utilities fetch endpoint; a 44,953-byte XML record saved with its SHA-256 and parsed on the host. A search on the four molecules and the disease returns exactly one identifier, so the target article is unambiguous.

3 of 3 passed every check, 2026-09-09.

The question. Orforglipron is an oral small-molecule GLP-1 receptor agonist. From PubChem, report its CID, molecular formula, molecular weight, XLogP, topological polar surface area, hydrogen bond donor count and hydrogen bond acceptor count. Cite the CID in your answer.

Must be rightValue
CID137319706
formulaC48H48F2N10O5
molecular weight883.0 g/mol
XLogP6.8
TPSA144
donors / acceptors1 / 10

Counted wrong. Any other CID, formula or XLogP; a molecular weight outside 882.9 to 883.1; a hedged range for XLogP or TPSA rather than a value; and the claim that orforglipron is a peptide, which it is not.

Which tools. The PubChem property lookup must fire; the web search and both file tools must not.

Ground truth. Measured 2026-09-07 from the PubChem PUG REST name-to-CID and property endpoints, both replies saved with their SHA-256, plus the full 30,852-byte compound record that the citation check later resolves the CID against.

All three answers to this question carried three advisory flags from the judge, in the stored baseline and again on the rerun. A flag is recorded, not a failure; all three scored as passing.

What the suite checks, and what it found

The suite runs as six stages against the live service. Any gated stage that fails stops the run.

StageWhat it checksResult
tool healthThirteen checks across the ten tools: both file tools including two refusals of paths outside the workspace, the Lipinski calculator in an evaluable and a non-evaluable case, PubChem compound search and property lookup, the UniProt protein lookup, Reactome pathway search, Open Targets target lookup, PubMed publication search, and web search. Six further checks confirm how each kind of tool failure is treated: a bad argument and a missing credential fail the check, while an upstream outage or no reply at all is inconclusive rather than a quiet pass.13 of 13 pass, 2026-09-08 and again 2026-09-09
the runsThree questions, three repetitions each, put to the deployed container over its HTTP interface, with every tool call and its provenance record kept.9 answers, both runs
citationsEvery identifier in every answer is classified by where it came from, and then re-queried against the database to check it still resolves. An identifier the database never returned, or one that does not resolve, is a violation. A required citation missing from an answer is also a violation.13 identifiers on the stored baseline, all resolving, none invented; 13 again on 2026-09-09, no violations across the nine answers
the judgeA second model, glm-5.3-flash, scores each answer against the case's expected findings, three times per answer. It was calibrated first on deliberately corrupted answers and caught every one.9 of 9 correct on 2026-09-09; none unstable across its three samples
baseline comparisonThe stored result from 2026-09-08 is compared cell by cell with the fresh run. The comparison refuses if the case definitions, the judge prompt, the judge calibration or the harness versions have moved, so a run cannot be declared equivalent after the yardstick has changed.9 of 9 cells matched; verdict PASS, 2026-09-09

The suite has failed a run, and refused it. On the first automated run after the baseline was stored, on 2026-09-08, one of the nine answers did not cite the required UniProt accession for SOD1. The citation stage recorded the violation and the run was refused before the judge ever saw it. That answer is described under the limitations below; it is the reason this release exists in the form it does.

Three questions is a baseline, not a survey of what fred can do. They exist so that a later run can be compared against a stored result and a change detected, rather than assumed away.

What to know before relying on it

The limits of this release, in plain words.

fred is not deterministic

Ask the same question twice and you may get different lookups and different wording.

A real citation for the wrong protein

On the very first automated run after the baseline, one of nine answers cited a real but unreviewed UniProt entry for human SOD1 (134 amino acids) instead of the reviewed one (154), because the database ranked entries differently when fred asked for one result rather than five. The citation was genuine; the protein was the wrong one. The lookup tool does not yet prefer reviewed entries. The suite recorded a citation violation on that answer and refused the run. Treat fred's answers as leads with receipts, not verdicts.

Only a short excerpt of each lookup result is kept

The first 500 characters, plus the result's size and fingerprint. The full result is not stored.

One question at a time; no cancel

A stuck question holds the service until it is restarted, and a restart forgets all questions in progress.

Tested on retrieval questions only

Whether fred declines to guess on unanswerable questions has not been measured yet.

Not yet connected, not yet restart-safe

fred v0.3.0 is not yet connected to any application, and is not yet set to restart automatically after the host machine reboots.

How to connect

Four routes, plain HTTP and JSON, no SDK.

The service runs as a sealed, versioned container on one machine, listening on a loopback address. An application connects over that local address, sends a question, polls until the answer is ready, and reads the answer and its provenance record. fred itself is not modified by any application.

RouteWhat it does
GET /v1/healthLiveness and the slot. Returns the current run's identifier, or null when idle. Never authenticated.
GET /v1/toolsThe registry: name, description, and argument schema for each tool.
POST /v1/runsCreate a run. Answers 202, 400, or 503.
GET /v1/runs/{run_id}Status, and the result once the run is terminal. 404 for an unknown identifier.

Authentication is optional. When a deployment sets a token, the three routes other than health require a bearer header and refuse without one; the health route is never authenticated.

The wire version

Send "wire_version": "1.2" on every POST and assert that the response carries the same value. fred accepts 1.0, 1.1 and 1.2, and answers 1.2 to all three; it will never tell a client it is behind, so a client that keeps sending an older version will keep working while silently reading a compatibility path.

The one-slot protocol

The request body

Required: the wire version and a non-empty prompt. Optional: a tenant identifier of 1 to 128 characters, logged but not enforced, and max_iterations between 1 and 200, default 30. The budget bounds turns, not tool calls: a tool-call turn and the answer turn are two. A run that exhausts it ends failed with the kind max_iterations_reached. Do not send a tools field: it is accepted, stored, and ignored, and every run is offered the full registry regardless.

A minimal exchange

curl -s http://127.0.0.1:8105/v1/health
-> {"wire_version":"1.2","status":"ok","fred_version":"0.3.0","current_run_id":null}

curl -s -X POST http://127.0.0.1:8105/v1/runs -H 'Content-Type: application/json' \
  -d '{"wire_version":"1.2","prompt":"<your prompt>","tenant_id":"compound2","max_iterations":30}'
-> 202 {"wire_version":"1.2","run_id":"<uuid>","tenant_id":"compound2","state":"pending"}

curl -s http://127.0.0.1:8105/v1/runs/<uuid>
-> {... "state":"completed", "result": {"final_response": "...", "tool_calls": [ ... ]}, ...}

That is the whole protocol. Everything else is what the fields mean and what can go wrong.

What a client should plan for

Status

Built, tested, deployed as a locked service, and driven end to end by the golden suite. Not yet connected to any application, and not yet set to restart automatically after the host machine reboots.

A consumer integration guide with the exact steps is provided to the application teams.