Tested models
Scores are from Pellaeon's own request sets, run through the real panel and checked by the program (the resulting ChimeraX state, the commands run, the cards shown). Local models ran on an RTX 4070 Super (12 GB). Last run: 2026-10-02, Pellaeon 0.2.8.
Which model
| Want | Pick | Notes |
|---|---|---|
| Highest score in the feature comparison | ChatGPT gpt-6-sol |
Plus/Pro subscription, sign in from Settings; about 10 s a request |
| Faster, same account | ChatGPT gpt-6-luna |
about 7 s a request |
| Free, nothing to install | Mistral ministral-8b-latest |
free key, about 3 s a request; vision |
| Free, alternative | Gemini gemini-flash-lite-latest |
free key, 15 requests a minute; vision; 32/32 on the precise-request set below |
| Private, on a 12 GB GPU | Ollama gemma4:12b |
runs on your computer |
| Private, faster | Ollama qwen3.5:9b |
about 2.5 s a request; a few more mistakes on display and ligand requests |
Scores
| Model | Everyday requests (152 × 3) | ChimeraX features (31 × 3) | Feature comparison (19 × 2) | Time per request |
|---|---|---|---|---|
gpt-6-sol |
35/38 | 11 s | ||
gpt-6-luna |
30/38 | 8 s | ||
ministral-8b-latest |
376/456 | 69/93 | 3 s | |
gemini-flash-lite-latest |
28/31 (one repeat) | 29/38 | 3 s | |
gemma4:12b |
383/456 | 80/93 | 21/38 | 5 s |
qwen3.5:9b |
357/456 | 2.5 s |
- Everyday requests: 152 requests mined from RBVI tutorials, the chimerax-users list and courses (open, colour, select, measure, label, compare, save). Used during development, so these scores are tuned.
- ChimeraX features: 31 requests covering the ChimeraX feature highlights (maps, morphs, glycans, symmetry, membranes, MD, ...); three repeats per model, one for Gemini.
- Feature comparison: 19 of the harder feature requests (map fitting, membrane orientation, palettes, worms, sequence, rotamers), run twice per model.
- Time is the median per request, from the same runs. A blank cell means that set was not run on that model.
- Differences of a few requests between runs are noise: two runs of the same code differ by up to 7 of 152.
Precise requests (32, Pellaeon 0.2.8, one repeat): chain scopes, gaps, insertion codes, signed numbering, conditionals, preservation. gpt-6-sol, gpt-6-luna, claude-sonnet-5.5 and gemini-flash-lite 32/32; gemma4:12b 31/32; claude-haiku-4.5 30/32; ministral-8b 28/32. Median response latency: Mistral 1.9 s, Gemini 2.1 s, Haiku 3.1 s, Sonnet 5.5 3.6 s, gemma 4.2 s, luna 5.1 s, sol 8.5 s. These 32 requests were used during development, so the scores measure regression, not an unseen benchmark.
Free-tier limits
As seen on new free accounts in September 2026; providers change them, and paid accounts differ.
| Provider | Free allowance | In practice |
|---|---|---|
| Mistral | about 190 requests a minute | no limit reached in any run |
| Gemini Flash-Lite | 15 a minute, plus a daily cap | enough for one person; a long run can hit the daily cap |
| Gemini Flash | 5 a minute | not usable: every miss was a rate limit |
| Groq | 7,000 to 8,000 input tokens a minute | not usable: about one request a minute |
OpenRouter openrouter/free |
50 requests a day | about a dozen questions |
One question is two to four requests of a few thousand tokens each (the state of your structures and the tool definitions travel with it).
Local models
Basic set of 38 requests, Pellaeon 0.2.0, September 2026. Larger did not mean better on this card: the 20B models spill out of 12 GB and slow down.
| Model | Basic set | Time | Notes |
|---|---|---|---|
qwen3:8b |
38/38 | 3 s | |
qwen3:4b |
37/38 | 20 s | 2.5 GB, the smallest that works |
gemma4:12b |
37/38 | 8 s | the default; vision |
gemma4:e4b |
37/38 | 5 s | vision, fits small cards |
granite4.1:8b |
37/38 | 5 s | |
qwen3:14b |
36/38 | 8 s | |
gpt-oss:20b |
36/38 | 15 s | spills out of 12 GB |
ministral-3:3b |
34/38 | 2 s | 3 GB; vision |
llama3.1:8b |
33/38 |
Below 30/38 and not recommended: command-r7b, cogito:8b, granite4:1b, qwen2.5:3b, granite4:micro, mistral-nemo:12b, lfm2.5:8b, llama3.2:3b, llama3-groq-tool-use:8b, granite3.3:2b, smollm2:1.7b, nemotron-mini:4b. No tool calling, so unusable: gemma3:12b, falcon3:3b, exaone3.5:2.4b.
Ollama keeps the previous model loaded for a few minutes; on a full card the new one runs on the CPU until the old one unloads.
Pellaeon