BenchmarksArchiveAgentContact

Theia

A model that reads fashion photographs.

Version oneOctober 2026Names the house, the season and the year

Benchmarks

How Theia compares

The share of photos each AI got right, on four tests. Longer bar is better. When an AI said “I don’t know”, that counted as wrong.

The Runway Test287 photos: house, season and year all right

Theia56.4%
Kimi K344.9%
Claude Opus 5.528.6%
GPT-6.1 Sol26.1%
Gemini 3.1 Pro23%
Claude Sonnet 5.521.6%
Kimi K2.612.2%
Qwen3.5-397B5.2%
Grok 4.71.4%

The Product House Test150 product photos: the house named

Theia100%
Claude Opus 5.598.7%
GPT-6.1 Sol95.3%
Claude Sonnet 5.592.7%
Gemini 3.1 Pro55.3%

The Blind Runway Check50 photos from kept-out shows: all three right

Theia56%
Gemini 3.1 Pro26%
Claude Opus 5.522%
Claude Sonnet 5.522%
GPT-6.1 Sol20%

The Exact Product Test75 products: the exact product named

Claude Opus 5.558.7%
GPT-6.1 Sol52%
Theia50.7%

Not every AI was run on every test, so some bars are missing. Theia leads three of the four tests. Claude Opus 5.5 leads exact product names.

Score and cost

Each dot is one AI on the runway photos. Up and to the left is better: more right, for less money. Pick a score to see how the picture changes.

Cost per photo in US dollars, on a log scale (each step is about 2 to 3 times more). Measured from the runs.

The model

Name
Theia, version one. The first model from ArchiveAgent.
Built from
Kimi K2.6, an open model from Moonshot AI.
How
Extra training (“fine-tuning”) on Tinker, a service from Thinking Machines. 5,025 examples, one pass, 80 minutes, $18.35.
You give it
One runway photo, or up to six photos of one product.
It gives back
The house (the brand), the season, the year, the product name, and how sure it is.
Not for
Chatting. It is a specialist. See “Where Theia falls short” below.

Mission

We are building an AI that understands fashion from the ground up.

Theia is the first step toward the ultimate ArchiveAgent: an AI designed for fashion.

  • TodayTheia names the house, the season and the year of a runway photo.
  • ThenLearn more houses, and look things up online to check its own answers.

What Theia gets right today, %

House87.8%
Season85.7%
Exact year64.1%
Exact product50.7%

House, season and year: 287 new runway photos. Exact product: 75 products, a different and harder test.

How we tested

Same photos. Same question.

Fresh photosTheia never trained on any test photo, show or product. We checked every file.

The same question for every AIEvery AI got the same photos and the same instructions, through OpenRouter. Theia ran on Tinker.

No credit for “I don’t know”A pass, or an answer that ran out of room, counts as wrong.

1

The Runway Test

287 runway photos from 41 shows that were kept out of Theia’s training. House, season and year all had to be right.

The question

Which house, season and year is this look from?

The right answer

Ann Demeulemeester, Fall/Winter, 2015

What each AI said

TheiaAnn Demeulemeester, Fall/Winter, 2015houseseasonyear
Claude Opus 5.5Ann Demeulemeester, Fall/Winter, no yearhouseseasonyear
GPT-6.1 SolAnn Demeulemeester, Fall/Winter, 2015houseseasonyear
Gemini 3.1 ProDior Homme, Fall/Winter, 2014houseseasonyear

All three parts must be right to score. “I don’t know” counts as wrong.

Photos fully right, %

Theia56.4%
Kimi K344.9%
Claude Opus 5.528.6%
GPT-6.1 Sol26.1%
Gemini 3.1 Pro23%

Why the others score low

Many of them said “I don’t know.” That counts as wrong. Theia was trained to always make its best guess.

Said “I don’t know” or ran out of room, %

Grok 4.760%
Kimi K2.645%
Claude Sonnet 5.540%
Gemini 3.1 Pro37%
Kimi K332%
Qwen3.5-397B29%
Claude Opus 5.525%
GPT-6.1 Sol9%
Theia0%

What the extra training added

12%to56%

Kimi K2.6, the AI we started from, then Theia after our extra lessons.

In shortTheia got 56% fully right. The next best, Kimi K3, got 45%. Claude Opus 5.5 got 29%.

2

The Product House Test

150 product photos from three houses: Louis Vuitton, Chanel and Rick Owens. The only job: name the house.

Three of the 150 product questions, one per house, picked at random (not by how well anyone did).

5 photos of this item. Tap one to look through them all.

The question

Here are photos of one item. Which house made it?

The right answer

Louis Vuitton

What each AI said

TheiaLouis Vuittonhouse
Claude Opus 5.5Louis Vuittonhouse
GPT-6.1 SolLouis Vuittonhouse
Gemini 3.1 Pro“I don’t know”house

A decline does not count. A collaboration, like “Louis Vuitton x Nike”, counts as naming the house.

Houses named right, %

Theia100%
Claude Opus 5.598.7%
GPT-6.1 Sol95.3%
Claude Sonnet 5.592.7%
Gemini 3.1 Pro55.3%

In shortTheia got all 150. Claude Opus 5.5 got 148. That is too close to call. Gemini 3.1 Pro, which often did not answer, got 55%.

3

The Blind Runway Check

A second, smaller runway test: 50 looks from shows kept completely out of training.

The question

Which house, season and year is this look from?

The right answer

Vetements, Spring/Summer, 2024

What each AI said

TheiaVetements, Spring/Summer, 2024houseseasonyear
Claude Opus 5.5Balenciaga, Fall/Winter, 2022houseseasonyear
GPT-6.1 SolBalenciaga, Spring/Summer, 2022houseseasonyear
Gemini 3.1 Pro“I don’t know”houseseasonyear

Same scoring as the Runway Test: all three must be right.

Photos fully right, %

Theia56%
Gemini 3.1 Pro26%
Claude Opus 5.522%
Claude Sonnet 5.522%
GPT-6.1 Sol20%

In shortTheia got 28 of 50. The best of the others got 13. With only 50 photos, treat this as a second opinion, not proof.

4

The Exact Product Test

The hardest one. Up to six photos of one product, and the AI must name it exactly, like “Christopher MM backpack.” A second AI judges whether the name matches.

Example 1 shows how the judge works. Examples 2 and 3 were picked at random from the 75 products.

6 photos of this item. Tap one to look through them all.

The question

Here are up to six photos of one item. What exactly is it?

The right answer

Louis Vuitton Christopher MM backpack

What each AI said

TheiaLouis Vuitton Christopher PM-style backpack in Damier Infini leathersame product
Claude Opus 5.5Louis Vuitton Christopher Backpack PMsame product
GPT-6.1 SolLouis Vuitton Christopher Backpacksame product

A second AI checks whether each answer is the same product as the right answer. Theia said “PM” where the answer is MM, so it is marked wrong.

Products named right, %

Claude Opus 5.558.7%
GPT-6.1 Sol52%
Theia50.7%

Why this one is hard

The name is not in the photoA brand can invent a name, like Gemini Bag. The AI has to have learned it.

Look-alikesA brand sells many similar bags and boots, each with its own name.

Close is wrongThe right kind of item with the wrong name does not score.

Right, by kind of name, % of products

Long descriptions7 or more words, like “black leather jacket with zip pockets” · 27 products

Theia70%
Claude Opus 5.574%
GPT-6.1 Sol70%

Short phrases4 to 6 words · 25 products

Theia40%
Claude Opus 5.564%
GPT-6.1 Sol56%

Made-up model names1 to 3 words, like “Gemini Bag” · 23 products

Theia39%
Claude Opus 5.535%
GPT-6.1 Sol26%

Every AI does fine when the right answer is a long description. The score drops when it is a short, made-up model name. That is the hardest slice: 23 of the 75 products. On it, Theia got 9 of 23, Claude Opus 5.5 8 and GPT-6.1 Sol 6. Two real cases:

6 photos of this item. Tap one to look through them all.

The question

Which exact product is this?

The right answer

Rick Owens Gemini Bag

What each AI said

TheiaRick Owens Gemini Bagsame product
Claude Opus 5.5Rick Owens Adri Bag (leather hobo with Cerberus-style clip hardware)same product
GPT-6.1 SolRick Owens Big Adri leather shoulder bagsame product

Rick Owens calls this bag “Gemini”. That is a name to learn, not something you can read off the photo.

6 photos of this item. Tap one to look through them all.

The question

Which exact product is this?

The right answer

Rick Owens Meganaska

What each AI said

TheiaRick Owens Cargotaconaskasame product
Claude Opus 5.5Rick Owens Bogun Wedge Bootssame product
GPT-6.1 SolRick Owens Zionic leather knee-high bootssame product

“Meganaska” and “Cargotaconaska” are two different Rick Owens boots in our data. Close is not enough.

A caution. Of those 23, 17 are from product lines that also appear in our training photos, and 6 are from lines Theia never saw. On the 17, Theia got 8, Opus 7 and Sol 5. On the 6, each got 1, 1 and 1. So part of Theia’s skill here is remembering lines it was shown. We have not yet tested with whole lines held out.

In shortThis is where Theia does not win. Claude Opus 5.5 named 59% of the products and Theia 51%. With only 75 products, a gap this size could be luck.

Who won what

Five contests

The Runway Testby 11.5 pointsTheia
The Product House Testby 2 of 150, too close to callTheia
The Blind Runway Checkby 15 of 50Theia
Hardest slice: made-up namesby 1 of 23, too close to callTheia
The Exact Product Testby 6 of 75Claude Opus 5.5
PointsTheia 4, Claude Opus 5.5 1

One point to the winner of each contest. It counts contests; it is not one grand score.

Cost

What it costs

Training Theia cost $18.35, one time. After that, naming a photo costs about 0.14 cents.

$18.35to train, once
$1.40to name 1,000 photos
9×cheaper per photo than Claude Opus 5.5

What you pay for each photo it gets completely right

Theia$0.0025
Kimi K3$0.0149
GPT-6.1 Sol$0.0195
Gemini 3.1 Pro$0.0265
Claude Sonnet 5.5$0.0273
Claude Opus 5.5$0.0434

In short. For each photo it gets fully right, Theia costs about 17 times less than Claude Opus 5.5. The $18.35 of training pays for itself after roughly 1,700 photos. We count the photos it gets wrong too, because you pay for those.

The build

How it was made

We did not build an AI from nothing. We took an open one, Kimi K2.6, and taught it fashion.

Build the testFresh photos Theia never trained on. We checked: none are shared.

Write the rules firstThe scoring rules were written down, dated and hashed before any run.

Same test for allOne script, the same photos, the same question, every AI through OpenRouter.

Teach it5,025 examples, one pass, 80 minutes.

3,000 runway2,025 products

Check our own workWe found three mistakes of our own, fixed them, and say so below.

ArchiveAgent and Theia are not affiliated with or endorsed by Moonshot AI or Thinking Machines. Kimi K2.6 is released under a modified MIT licence.

Honest notes

Where Theia falls short

  • It always answers, even when unsure. It also says how sure it is: keep its most confident 75% of answers and the house is right 98.6% of the time.
  • It knows 14 houses well. On photos from houses it never trained on, it names the house 57% of the time, against 90% on the others.
  • Exact product names are its weak spot (Test 4).
  • On studio product photos, the season and year are mostly guesses.
  • One run per AI, scored by us. We found and fixed three mistakes in our own scoring: it ignored pre-fall, it let a few product lines leak between training and test, and it marked “Louis Vuitton x Nike” as the wrong house.
  • It is a specialist, not a chat assistant.

At a glance

Theia in six numbers

56%new runway photos with house, season and year all right
88%new runway photos with the house right
150/150product photos with the house right
17×cheaper than Claude Opus 5.5 per photo fully right
$18.35to train, in 80 minutes
4 of 5contests won. Exact product names went to Claude Opus 5.5.
For researchers: every number in one table
TheiaClaude Opus 5.5GPT-6.1 SolGemini 3.1 ProClaude Sonnet 5.5Grok 4.7Kimi K3Kimi K2.6Qwen3.5-397B
Runway: house287 photos87.8%74.2%66.2%49.5%59.2%32.1%61.0%52.6%33.8%
Runway: season287 photos85.7%60.6%56.4%52.6%47.0%8.7%59.2%16.7%48.8%
Runway: year287 photos64.1%30.3%30.0%30.0%23.0%1.4%48.1%12.2%18.1%
Runway: all three287 photos56.4%28.6%26.1%23.0%21.6%1.4%44.9%12.2%5.2%
Product photos: house150 cases100.0%98.7%95.3%55.3%92.7%20.7%———
Blind runway: all three50 photos56.0%22.0%20.0%26.0%22.0%0.0%———
Exact product75 products50.7%58.7%52.0%——————
Cost per photoUSD, measured$0.0014$0.0124$0.0051$0.0061$0.0059$0.0075$0.0067$0.0005$0.0006

Runway rows: 287 new photos from 41 shows; house, season and year scored exactly (pre-fall and resort count as their own seasons). Product house: 150 cases; a collaboration answer such as “Louis Vuitton x Nike” counts as naming the house. Blind runway: 50 photos. Exact product: 75 products, same prompt for every AI, a second AI judges; only three systems were run on it. Every comparison AI ran through OpenRouter with the same instructions; Theia ran on Tinker. Those instructions say abstaining is better than guessing, which some AIs follow more than others; a decline counts as wrong. Costs are measured per photo from the runs. One run per system; labels are model-made and trusted by the owner.