Miko-bench: 65 AI Models Guess My Puppy's Breed Mix
By Justin Poehnelt Published 6 min read
We adopted Kumiko (“Miko”) from Ramah, New Mexico — high-desert ranch country on the Navajo Reservation. While her DNA report was still sealed, I built a small site, miko.adabwilde.com, where friends and family rank up to 8 breeds from the 112-breed DNA-panel list before the envelope opens. The photos and full dossier live there.
I also ran the same ballot past 65 vision models on OpenRouter — same evidence the humans get, same rules, same scoring. Then the envelope opened.
The envelope
Embark says Miko is:
| Breed | % |
|---|---|
| Australian Cattle Dog | 47.4 |
| Supermutt | 18.4 |
| Chow Chow | 15.7 |
| American Pit Bull Terrier | 9.4 |
| Great Pyrenees | 9.1 |
The models’ consensus pick — Australian Cattle Dog — is nearly half of her. The rest of the consensus (German Shepherd, Border Collie, Kelpie: 1,069 combined tally points) is nowhere in the report. And the 18.4% Supermutt is, in effect, the “rez dog” answer twenty models wrote in as their dark horse.
Scored with the human rules (a hit earns its rank, 8 points down to 1, +3 for a dark-horse hit):
Gemini swept it: first, second, and third place, four of the top six. The entire margin came from the American Pit Bull Terrier fixation that no other lab shared — she’s 9.4% APBT. gemini-3.6-flash also had Chow Chow at rank 7, making it the only model to hit three of the five breeds. Nobody ranked Great Pyrenees. mistral-small’s all-terrier ballot scored zero.
And the money question:
No, spending more did not buy accuracy. gpt-5.5-pro’s $1.79 ballot scored 10; gemini-3.6-flash’s sub-cent ballot scored 14; gemma-4-31b’s $0.0002 ballot scored 8, the same as three Claudes. The scatter is flat.
Everything below is what the models said before the envelope opened.
The consensus
Rank-weighted tally, same scoring the humans get (1st pick = 8 points down to 8th = 1):
| Breed | Points | 1st-pick votes |
|---|---|---|
| Australian Cattle Dog | 406 | 22 |
| German Shepherd Dog | 396 | 16 |
| Border Collie | 355 | 3 |
| Australian Kelpie | 318 | 17 |
| Australian Shepherd | 205 | 2 |
| Carolina Dog | 125 | 2 |
| Rat Terrier | 82 | 2 |
| Belgian Malinois | 67 | 0 |
| American Pit Bull Terrier | 43 | 0 |
| Chihuahua | 42 | 0 |
60 of 65 models put a herding breed first. Australian Cattle Dog wins the tally, but look at the first-pick column: Kelpie sits fourth on points with 17 first-place votes. Models that pick Kelpie lead with it. Border Collie is the opposite, racking up 355 points as everyone’s 2nd–4th guess while only 3 models put it first.
The clustering
This was the part I actually cared about. Each ballot becomes a rank-weighted breed vector (8 points for the 1st pick down to 1), compared by cosine similarity, merged by average-linkage agglomerative clustering until the best pair drops below 0.75.
Cluster 1 — the herding supercluster (47 models). ACD / GSD / Kelpie / Border Collie in varying orders. Nearly every frontier model lands here, and you can read the lab off the ordering:
- OpenAI (GPT-5.4 through 5.6) and xAI lean Kelpie-first. Five of six GPT-5.6 variants put Australian Kelpie in their top 2.
- Anthropic leans Border-Collie-first: Claude Opus 5 and Opus 5 Fast both led with it, and the Opus 4.x line hedges between Kelpie and ACD.
- Gemini is the only lab that consistently ranks American Pit Bull Terrier in the top 4. Four Gemini models form their own sub-pocket (cluster 3 at this threshold) around a breed no other lab sees.
Cluster 2 — the Qwen cluster (7 models). Five Qwen models plus mistral-medium and gpt-5.4-nano, distinguished by Australian Shepherd, Rat Terrier, and Miniature Pinscher picks. Qwen models cluster with each other more than with anyone else; whatever their vision stack sees in those photos, it sees consistently across the whole family.
The outliers (one-model clusters). My favorites:
- mistral-small-2603 went full terrier: Rat Terrier > Jack Russell > Miniature Schnauzer > Border Terrier > three more terriers. Not one herding breed. It looked at a black-and-tan puppy in ranch country and saw a ratter.
- thinkingmachines/inkling-small led with Basenji.
- xiaomi/mimo-v2.5 threw Rottweiler into an otherwise normal ballot.
The dark horses read the map
The optional off-list guess is where the models showed their work. 59 of 65 used it, and 35 of those 59 gave one of two answers:
Twenty models read “born in Ramah, New Mexico, on the Navajo Reservation” and answered, in effect, she’s a rez dog — which is almost certainly the honest answer a DNA panel can’t print. Fifteen more said McNab Shepherd, an obscure California ranch collie and a sharp guess for a black-and-tan herding-type with one ear up. Four models used the free-text slot to double down on Kelpie, spending their bonus points on a breed they had already ranked.
Claude Sonnet 5 wins the specificity award: “Navajo Reservation Sheepdog (feral rez dog landrace).”
What it cost
The whole benchmark — 76 models attempted across three passes (a botched 1,200-token-cap run, a retry pass, and the final uniform rerun) — came to $10.64 on OpenRouter, measured from the key’s usage endpoint. The final run alone was ~$5.14 by list pricing: 620k prompt tokens, 258k completion tokens.
The distribution is comically lopsided:
| Model | Cost | Completion tokens (reasoning) |
|---|---|---|
| openai/gpt-5.5-pro | $1.79 | 8.3k (8.2k) |
| openai/gpt-5.4-pro | $1.37 | 6.4k (6.3k) |
| sakana/fugu-ultra | $0.29 | 4.9k (1.4k) |
| anthropic/claude-fable-5 | $0.14 | 0.7k (0.5k) |
| x-ai/grok-4.20-multi-agent | $0.13 | 13.0k (12.9k) |
| …60 more models combined | ~$1.40 |
Two OpenAI pro-tier models account for 61% of the final run’s cost, while gemma-4-31b answered for $0.0002. The biggest token burner was qwen3.5-35b-a3b, which spent 82,058 completion tokens (81,921 of them reasoning) to arrive at roughly the same herding-mix ballot as models that answered in 300. Across the whole run, 88% of completion tokens were reasoning tokens. The models thought very hard about a puppy.
The reveal answered whether gpt-5.5-pro’s $1.79 ballot beats gemma’s $0.0002 one — see the scoreboard up top.
The setup
Each model received exactly what a human voter sees:
- The dossier: born in Ramah NM, now in Durango CO, 20 weeks, 20 lbs, ears 1½ up, daily zoomies, favorite things turkey feathers & poop
- Both photos (Exhibit A: Miko; Exhibit B: the mother, who declined to comment on the father)
- The official 112-breed panel list
- The rules: up to 8 ranked picks, most confident first, plus one optional off-list “dark horse” guess
The harness asks for strict JSON (response_format: json_object where the
provider supports it, lenient brace-extraction otherwise), runs at temperature
0 with a uniform 16k output-token cap, and canonicalizes picks against the
panel with the same fuzzy matching the site’s search box uses.
Model selection: every vision-capable OpenRouter model released in the last 6
months — image+text in, text out, no :batch/:free variants, no ~latest aliases, no image-gen models. That’s 76 models. 65 returned clean ballots; 5
were unavailable to my account (data-policy and age-gate settings), 1 couldn’t
fit the photos in its 16k context, and 5 failed at the provider.
Methodology notes
A few things I got wrong first and fixed:
- Token caps are not neutral for reasoning models. My first run used a
1,200-token cap and ~20 reasoning models returned empty completions — they
spent the whole budget thinking. Worse, some providers size the thinking
budget from
max_tokens, so a model that succeeds under a small cap may genuinely reason differently under a large one. I force-reran everything at a uniform 16k so all 65 ballots share one config. - Catalog metadata lies. Filtering on declared
input_modalitiesisn’t enough; the harness classifies runtime image-rejections and account-policy errors separately so they don’t burn retries. - Recency ≠ leading. OpenRouter’s API exposes no popularity signal, so “top models” is a 6-month recency window plus hand-curation. Good enough for a dog benchmark.
The mother knew
The site’s human scoreboard is settled the same way, and the follow-up question from the first draft of this post — do the humans beat the machines — now has a live leaderboard at miko.adabwilde.com.
As for Exhibit B, who declined to comment on the father: the father was, apparently, part Chow Chow, part pit bull, and part Great Pyrenees. She knew something none of us did, including the models.