Which model

The measured list, by VRAM

ResearchZosho ships no model. Research runs need a model server that speaks the OpenAI chat API. That can be one on your own machine (llama.cpp, Ollama, LM Studio) or a hosted API with a key. Asking what the library already holds works without a model.

VRAM means the memory on your graphics card, not the computer’s RAM. nvidia-smi shows it on Linux and Windows. On a Mac with Apple silicon the card shares the machine’s unified memory, so read the tiers against about two thirds of that.

Eleven models were run through the same research questions on the same library, on one machine with one card. Each was judged on what came out: whether the facts were right, how many claims the write-up made, how many of its citations the citation checker could read against the source, and how long a question took. The 9B and the 27B were then compared on five more paired questions. Both got the facts right. The 27B made 25 claims to the 9B's 10, and 26 checkable citations to the 9B's 1. Everything below comes from those runs. A model not on the list was not measured. That is different from not recommended.

researchzosho models prints the row for the card it finds, with the command that serves it. researchzosho model install sets up the top row of that tier on demand. The model comes up when a run needs it and goes away after twenty idle minutes. On a Mac it uses llama.cpp's Metal build. On Windows it uses the Vulkan build, which runs on any card.

24 GB of VRAM or more

ModelFileWhat we saw
Qwen3.8-27B at 4-bitQwen3.8-27B-UD-Q4_K_M.ggufthe reference: the deepest answers, the most claims, the most citations the checker can read; 17 GB file

16 GB of VRAM

ModelFileWhat we saw
gpt-oss-20bgpt-oss-20b-F16.ggufright, 8 claims, about 5 minutes a question; 13 GB in use with two 16k slots. The choice when speed matters
Gemma 4 26B-A4B at 4-bitgemma-4-26B-A4B-it-UD-IQ4_XS.ggufright and the most careful writer, 10 claims, about nine times slower; 14.6 GB in use
Gemma 4 12B at 4-bitgemma-4-12b-it-Q4_K_M.ggufright, 10 claims, six times slower; 9 GB in use, the most room for context

8 GB of VRAM

ModelFileWhat we saw
Qwen3.5 9B at 4-bitQwen3.5-9B-Q4_K_M.ggufright and deep, 10 claims, half its citations readable, about 11 minutes a question; 5.9 GB in use with one 16k slot
Gemma 4 12B, smaller 4-bit filegemma-4-12b-it-IQ4_XS.ggufright, 10 claims, about 30 minutes a question; 7.3 GB in use with one 16k slot

4 GB of VRAM

ModelFileWhat we saw
Gemma 4 E4B at 4-bitgemma-4-E4B-it-Q4_K_M.ggufright, 9 claims, the best citation reader of the small models, about 6 minutes a question; 3.6 GB in use

2 GB of VRAM

ModelFileWhat we saw
Gemma 4 E2B at 4-bitgemma-4-E2B-it-Q4_K_M.ggufthe claims came out right, but its write-ups were headings with no text; 2 GB in use. A hosted API is the better answer this small

The files are the ones Hugging Face lists under unsloth/<model>-GGUF. model install checks every download against a recorded sha256 and refuses a mismatch.

Measured and not recommended

What the runs need from a server

The chat

researchzosho chat runs on the same model as the runs. The same two-turn conversation was tried on five models against one library. The first turn asked what the shelves hold on a subject. The second asked which entry is the most recent and who wrote it.

The 9B and gpt-oss-20b hold the conversation. The 27B sees more. The 12B answers but does not dig. The 4B cannot be trusted to cite. The chat's tool list is about 3,900 tokens, so 16k of context or more leaves room for the entries it opens.

A hosted API instead

Any OpenAI-compatible endpoint works. You give the base address, the model name the provider expects, and the key. Your documents go to that server and nowhere else. researchzosho setup asks for these and checks them with one question. A second server can do the judging while a local model reads. The full page and the guide's section 12 have that.