Which model
The measured list, by VRAM
ResearchZosho ships no model. Research runs need a model server that speaks the OpenAI chat API. That can be one on your own machine (llama.cpp, Ollama, LM Studio) or a hosted API with a key. Asking what the library already holds works without a model.
VRAM means the memory on your graphics card, not the computer’s RAM. nvidia-smi shows it
on Linux and Windows. On a Mac with Apple silicon the card shares the machine’s unified memory, so read the tiers
against about two thirds of that.
Eleven models were run through the same research questions on the same library, on one machine with one card. Each was judged on what came out: whether the facts were right, how many claims the write-up made, how many of its citations the citation checker could read against the source, and how long a question took. The 9B and the 27B were then compared on five more paired questions. Both got the facts right. The 27B made 25 claims to the 9B's 10, and 26 checkable citations to the 9B's 1. Everything below comes from those runs. A model not on the list was not measured. That is different from not recommended.
researchzosho models prints the row for the card it finds, with the command that serves it.
researchzosho model install sets up the top row of that tier on demand. The model comes up when a
run needs it and goes away after twenty idle minutes. On a Mac it uses llama.cpp's Metal build. On Windows it
uses the Vulkan build, which runs on any card.
24 GB of VRAM or more
| Model | File | What we saw |
|---|---|---|
| Qwen3.8-27B at 4-bit | Qwen3.8-27B-UD-Q4_K_M.gguf | the reference: the deepest answers, the most claims, the most citations the checker can read; 17 GB file |
16 GB of VRAM
| Model | File | What we saw |
|---|---|---|
| gpt-oss-20b | gpt-oss-20b-F16.gguf | right, 8 claims, about 5 minutes a question; 13 GB in use with two 16k slots. The choice when speed matters |
| Gemma 4 26B-A4B at 4-bit | gemma-4-26B-A4B-it-UD-IQ4_XS.gguf | right and the most careful writer, 10 claims, about nine times slower; 14.6 GB in use |
| Gemma 4 12B at 4-bit | gemma-4-12b-it-Q4_K_M.gguf | right, 10 claims, six times slower; 9 GB in use, the most room for context |
8 GB of VRAM
| Model | File | What we saw |
|---|---|---|
| Qwen3.5 9B at 4-bit | Qwen3.5-9B-Q4_K_M.gguf | right and deep, 10 claims, half its citations readable, about 11 minutes a question; 5.9 GB in use with one 16k slot |
| Gemma 4 12B, smaller 4-bit file | gemma-4-12b-it-IQ4_XS.gguf | right, 10 claims, about 30 minutes a question; 7.3 GB in use with one 16k slot |
4 GB of VRAM
| Model | File | What we saw |
|---|---|---|
| Gemma 4 E4B at 4-bit | gemma-4-E4B-it-Q4_K_M.gguf | right, 9 claims, the best citation reader of the small models, about 6 minutes a question; 3.6 GB in use |
2 GB of VRAM
| Model | File | What we saw |
|---|---|---|
| Gemma 4 E2B at 4-bit | gemma-4-E2B-it-Q4_K_M.gguf | the claims came out right, but its write-ups were headings with no text; 2 GB in use. A hosted API is the better answer this small |
The files are the ones Hugging Face lists under unsloth/<model>-GGUF. model install
checks every download against a recorded sha256 and refuses a mismatch.
Measured and not recommended
- Qwen3 14B and Mistral Small 24B: a fact wrong each; Mistral needs 19 GB.
- The 27B at 3-bit: slow, and it under-finds.
- Qwen3 8B and Llama 3.1 8B: thin answers.
- Granite 4.1 8B: right, but 9.4 GB in use with one 16k slot, so not an 8 GB fit, and one unit slip.
- Phi-4-reasoning-plus, Falcon-H1 7B, SmolLM3 3B, LFM2 8B, Granite 4.0 micro: copied the prompt, answered without sources, or got the facts wrong.
What the runs need from a server
- Tool calling. The readers search and fetch through tools. With llama.cpp that means
--jinja, so the server applies the model's own chat template. Without it, tool calls come back as prose. - Context. 16k per reader is the floor the list was measured at. 32k gives the write-up more
evidence. The
--parallelcount is how many readers work at once. - Reasoning off or low where the model has the switch. Otherwise a reader spends its turn
thinking instead of reading. The serve commands from
researchzosho modelscarry the right setting per row.
The chat
researchzosho chat runs on the same model as the runs. The same two-turn conversation was tried
on five models against one library. The first turn asked what the shelves hold on a subject. The second asked
which entry is the most recent and who wrote it.
- Qwen 3.8 27B: found all six relevant entries, sorted and cited them, answered the follow-up from one look-up. 27 s and 9 s a turn.
- gpt-oss-20b: found two of the three books, cited them, opened the entries for the follow-up and answered right. 4 s a turn.
- Gemma 4 12B: found and cited the entries, but did not open them for the follow-up and said it did not know dates the entries carry.
- Qwen 3.5 9B: found two of three, cited them, opened both for the follow-up and answered right. 4 s a turn.
- Qwen 3.5 4B: right on one run; on another it cited four entry ids that do not exist. The chat now admits such a reference under the reply.
The 9B and gpt-oss-20b hold the conversation. The 27B sees more. The 12B answers but does not dig. The 4B cannot be trusted to cite. The chat's tool list is about 3,900 tokens, so 16k of context or more leaves room for the entries it opens.
A hosted API instead
Any OpenAI-compatible endpoint works. You give the base address, the model name the provider expects, and
the key. Your documents go to that server and nowhere else. researchzosho setup asks for these and
checks them with one question. A second server can do the judging while a local model reads. The
full page and the guide's
section 12 have that.