Miko-bench: 65 AI Models Guess My Puppy's Breed Mix
We adopted Kumiko (“Miko”) from Ramah, New Mexico — high-desert ranch country on the Navajo Reservation. While her DNA report was still sealed, I built a small site, miko.adabwilde.com, where friends and family rank up to 8 breeds from the 112-breed DNA-panel list before the envelope opens. The photos and full dossier live there.
I also ran the same ballot past 65 vision models on OpenRouter — same evidence the humans get, same rules, same scoring. Then the envelope opened.
The envelope
Embark says Miko is:
| Breed | % |
|---|---|
| Australian Cattle Dog | 47.4 |
| Supermutt | 18.4 |
| Chow Chow | 15.7 |
| American Pit Bull Terrier | 9.4 |
| Great Pyrenees | 9.1 |
The models’ consensus pick — Australian Cattle Dog — is nearly half of her. The rest of the consensus (German Shepherd, Border Collie, Kelpie: 1,069 combined tally points) is nowhere in the report. And the 18.4% Supermutt is, in effect, the “rez dog” answer twenty models wrote in as their dark horse.
Scored with the human rules (a hit earns its rank, 8 points down to 1, +3 for a dark-horse hit):
Bar chart of the top 15 accuracy scores out of 65 models, colored by lab. gemini-3.6-flash leads with 14 of 30, then gemini-3.7-flash at 13, gemini-3.1-flash-lite, gpt-chat-latest, and ox-alpha at 12, followed by more Gemini and GPT models at 11 and a row of models tied at 8.
Gemini swept it: first, second, and third place, four of the top six. The entire margin came from the American Pit Bull Terrier fixation that no other lab shared — she’s 9.4% APBT. gemini-3.6-flash also had Chow Chow at rank 7, making it the only model to hit three of the five breeds. Nobody ranked Great Pyrenees. mistral-small’s all-terrier ballot scored zero.
And the money question:
Scatter plot of accuracy score versus ballot cost on a log axis, colored by lab. Scores show no upward trend with cost: gemini-3.6-flash tops the chart at under a cent, gemma-4-31b scores 8 at two hundredths of a cent, gpt-5.5-pro scores 10 at $1.79, and mistral-small sits at zero.
No, spending more did not buy accuracy. gpt-5.5-pro’s $1.79 ballot scored 10; gemini-3.6-flash’s sub-cent ballot scored 14; gemma-4-31b’s $0.0002 ballot scored 8, the same as three Claudes. The scatter is flat.
Everything below is what the models said before the envelope opened.
The consensus
Rank-weighted tally, same scoring the humans get (1st pick = 8 points down to 8th = 1):
Horizontal bar chart of the rank-weighted tally across 65 model ballots: Australian Cattle Dog 406 points with 22 first picks, German Shepherd Dog 396 with 16, Border Collie 355 with 3, Australian Kelpie 318 with 17, then Australian Shepherd, Carolina Dog, Rat Terrier, Belgian Malinois, American Pit Bull Terrier, and Chihuahua.
| Breed | Points | 1st-pick votes |
|---|---|---|
| Australian Cattle Dog | 406 | 22 |
| German Shepherd Dog | 396 | 16 |
| Border Collie | 355 | 3 |
| Australian Kelpie | 318 | 17 |
| Australian Shepherd | 205 | 2 |
| Carolina Dog | 125 | 2 |
| Rat Terrier | 82 | 2 |
| Belgian Malinois | 67 | 0 |
| American Pit Bull Terrier | 43 | 0 |
| Chihuahua | 42 | 0 |
60 of 65 models put a herding breed first. Australian Cattle Dog wins the tally, but look at the first-pick column: Kelpie sits fourth on points with 17 first-place votes. Models that pick Kelpie lead with it. Border Collie is the opposite, racking up 355 points as everyone’s 2nd–4th guess while only 3 models put it first.
Heatmap of how many models ranked each of the top eight breeds at each ballot position. Australian Cattle Dog, German Shepherd Dog, and Australian Kelpie concentrate in ranks 1 and 2; Border Collie peaks at rank 3 with 22 models; Australian Shepherd peaks at rank 5; Carolina Dog and Rat Terrier skew to the back of the ballot.
The clustering
This was the part I actually cared about. Each ballot becomes a rank-weighted breed vector (8 points for the 1st pick down to 1), compared by cosine similarity, merged by average-linkage agglomerative clustering until the best pair drops below 0.75.
Scatter plot of 64 ballots projected to two dimensions by classical MDS of cosine distance, colored by lab. A dense central mass holds OpenAI, Anthropic, xAI, and Google; Qwen points scatter toward the upper left; labeled outliers include inkling-small (Basenji), mimo (Rottweiler), and qwen3.5-122b (Rat Terrier). mistral-small's all-terrier ballot is excluded because it shares nothing with any other ballot.
Cluster 1 — the herding supercluster (47 models). ACD / GSD / Kelpie / Border Collie in varying orders. Nearly every frontier model lands here, and you can read the lab off the ordering:
- OpenAI (GPT-5.4 through 5.6) and xAI lean Kelpie-first. Five of six GPT-5.6 variants put Australian Kelpie in their top 2.
- Anthropic leans Border-Collie-first: Claude Opus 5 and Opus 5 Fast both led with it, and the Opus 4.x line hedges between Kelpie and ACD.
- Gemini is the only lab that consistently ranks American Pit Bull Terrier in the top 4. Four Gemini models form their own sub-pocket (cluster 3 at this threshold) around a breed no other lab sees.
Dot plot of the average rank each lab gave six breeds, where 9 means left off the ballot. OpenAI and xAI average about rank 2 on Australian Kelpie while Google and Qwen average near 7; Google averages about rank 6 on American Pit Bull Terrier while every other lab leaves it near 9; Anthropic ranks Border Collie highest of any lab.
Cluster 2 — the Qwen cluster (7 models). Five Qwen models plus mistral-medium and gpt-5.4-nano, distinguished by Australian Shepherd, Rat Terrier, and Miniature Pinscher picks. Qwen models cluster with each other more than with anyone else; whatever their vision stack sees in those photos, it sees consistently across the whole family.
The outliers (one-model clusters). My favorites:
- mistral-small-2603 went full terrier: Rat Terrier > Jack Russell > Miniature Schnauzer > Border Terrier > three more terriers. Not one herding breed. It looked at a black-and-tan puppy in ranch country and saw a ratter.
- thinkingmachines/inkling-small led with Basenji.
- xiaomi/mimo-v2.5 threw Rottweiler into an otherwise normal ballot.
The dark horses read the map
The optional off-list guess is where the models showed their work. 59 of 65 used it, and 35 of those 59 gave one of two answers:
Bar chart of dark-horse guesses across 65 ballots: rez or village dog 20, one-off guesses 16, McNab Shepherd 15, no dark horse 6, Miniature American Shepherd 4, Australian Kelpie smuggled back on-list 4.
Twenty models read “born in Ramah, New Mexico, on the Navajo Reservation” and answered, in effect, she’s a rez dog — which is almost certainly the honest answer a DNA panel can’t print. Fifteen more said McNab Shepherd, an obscure California ranch collie and a sharp guess for a black-and-tan herding-type with one ear up. Four models used the free-text slot to double down on Kelpie, spending their bonus points on a breed they had already ranked.
Claude Sonnet 5 wins the specificity award: “Navajo Reservation Sheepdog (feral rez dog landrace).”
What it cost
The whole benchmark — 76 models attempted across three passes (a botched 1,200-token-cap run, a retry pass, and the final uniform rerun) — came to $10.64 on OpenRouter, measured from the key’s usage endpoint. The final run alone was ~$5.14 by list pricing: 620k prompt tokens, 258k completion tokens.
The distribution is comically lopsided:
Log-log scatter of ballot cost versus completion tokens, colored by lab. gemma-4-31b sits at the cheap low-token corner, qwen3.5-35b at 82k tokens for about ten cents, grok-4.20-multi-agent at 13k tokens, and gpt-5.5-pro and gpt-5.4-pro alone past the one-dollar line.
| Model | Cost | Completion tokens (reasoning) |
|---|---|---|
| openai/gpt-5.5-pro | $1.79 | 8.3k (8.2k) |
| openai/gpt-5.4-pro | $1.37 | 6.4k (6.3k) |
| sakana/fugu-ultra | $0.29 | 4.9k (1.4k) |
| anthropic/claude-fable-5 | $0.14 | 0.7k (0.5k) |
| x-ai/grok-4.20-multi-agent | $0.13 | 13.0k (12.9k) |
| …60 more models combined | ~$1.40 |
Two OpenAI pro-tier models account for 61% of the final run’s cost, while gemma-4-31b answered for $0.0002. The biggest token burner was qwen3.5-35b-a3b, which spent 82,058 completion tokens (81,921 of them reasoning) to arrive at roughly the same herding-mix ballot as models that answered in 300. Across the whole run, 88% of completion tokens were reasoning tokens. The models thought very hard about a puppy.
The reveal answered whether gpt-5.5-pro’s $1.79 ballot beats gemma’s $0.0002 one — see the scoreboard up top.
The setup
Each model received exactly what a human voter sees:
- The dossier: born in Ramah NM, now in Durango CO, 20 weeks, 20 lbs, ears 1½ up, daily zoomies, favorite things turkey feathers & poop
- Both photos (Exhibit A: Miko; Exhibit B: the mother, who declined to comment on the father)
- The official 112-breed panel list
- The rules: up to 8 ranked picks, most confident first, plus one optional off-list “dark horse” guess
The harness asks for strict JSON (response_format: json_object where the
provider supports it, lenient brace-extraction otherwise), runs at temperature
0 with a uniform 16k output-token cap, and canonicalizes picks against the
panel with the same fuzzy matching the site’s search box uses.
Model selection: every vision-capable OpenRouter model released in the last 6
months — image+text in, text out, no :batch/:free variants, no ~latest aliases, no image-gen models. That’s 76 models. 65 returned clean ballots; 5
were unavailable to my account (data-policy and age-gate settings), 1 couldn’t
fit the photos in its 16k context, and 5 failed at the provider.
Methodology notes
A few things I got wrong first and fixed:
- Token caps are not neutral for reasoning models. My first run used a
1,200-token cap and ~20 reasoning models returned empty completions — they
spent the whole budget thinking. Worse, some providers size the thinking
budget from
max_tokens, so a model that succeeds under a small cap may genuinely reason differently under a large one. I force-reran everything at a uniform 16k so all 65 ballots share one config. - Catalog metadata lies. Filtering on declared
input_modalitiesisn’t enough; the harness classifies runtime image-rejections and account-policy errors separately so they don’t burn retries. - Recency ≠ leading. OpenRouter’s API exposes no popularity signal, so “top models” is a 6-month recency window plus hand-curation. Good enough for a dog benchmark.
The mother knew
The site’s human scoreboard is settled the same way, and the follow-up question from the first draft of this post — do the humans beat the machines — now has a live leaderboard at miko.adabwilde.com.
As for Exhibit B, who declined to comment on the father: the father was, apparently, part Chow Chow, part pit bull, and part Great Pyrenees. She knew something none of us did, including the models.
Opinions are my own and not the views of my employer.
© 2026 by Justin Poehnelt is licensed under CC BY-SA 4.0