Benchmark Suite · 7,432 Core Entities · 4 Domains

AgriTaxon

Diagnosing the Long-Tail Agricultural Knowledge Gap of Large Multimodal Models

Xin Zeng, Benfeng Xu, Qian Chen, Jialin Kuai, Wentao Zhang, Liguo Lang, Shancheng Fang, Huarui Wu

Under Review

1,971 Crops 178 Livestock 3,485 Pests 1,798 Weeds

Open-Ended Taxonomic Naming for Agricultural AI

Open-ended taxonomic naming is a foundational capability for intelligent agricultural decision-making. It requires a model to generate a standardized name for an organism or breed in an image without relying on a predefined label set. Most agricultural vision benchmarks instead assume a fixed label set, leaving unclear whether Large Multimodal Models (LMMs) can independently produce the correct name.

AgriTaxon is a knowledge-grounded and AI-ready benchmark suite spanning Crops, Livestock, Pests, and Weeds. Each core entity is linked through Wikidata to FAO Ecocrop, FAO DAD-IS, or the EPPO Global Database and is paired with machine-readable identifiers, a canonical name, collected aliases, a domain label, and an executable evaluation protocol. We evaluate 14 LMMs using four-option recognition and open-ended naming. Alias validation by GPT-5 Mini agrees with a domain expert on 98% of 100 sampled exact-match errors.

7,432
Core Entities
1,052
Hard Cases
14,845
Wild Images
14
LMMs Evaluated

Core, Hard, and In-the-Wild Evaluation

The core benchmark measures broad agricultural entity naming. AgriTaxon-Hard is a performance-conditioned diagnostic subset, while AgriTaxon-Wild evaluates the same naming capability using multiple independent field observations.

DatasetPurposeEntitiesImages
AgriTaxon CoreBroad taxonomic-naming benchmark7,4327,432
AgriTaxon-HardPerformance-conditioned diagnostic stress test1,0521,052
AgriTaxon-WildIn-the-wild multi-image evaluation2,96914,845

AgriTaxon-Hard contains core samples answered correctly by at most two of the 14 evaluated models under the default multiple-choice protocol. It is intended for failure analysis and tool-assisted recovery, not as an estimate of average benchmark performance.

AgriTaxon-Wild uses strict Wikidata P3151 alignment to iNaturalist, without fuzzy name matching. It retains 2,969 entities with five qualifying field photographs each and evaluates one, three, or five views of the same entity.

Leaderboard

Accuracy (%) on AgriTaxon for 14 LMMs, ranked by OE-Acc within each block. MC Macro, OE-EM, and OE-Acc are unweighted macro-averages across the four tracks; the four track-specific open-ended columns report EM. Bold = best per column.

Model Release Hard Multi-Choice Open-Ended
Crop Live. Pest Weed Macro Crop Live. Pest Weed EM Acc
Proprietary Models
gemini-3-pro-preview 2025.11 9.0 83.795.572.578.182.5 43.649.430.053.044.051.2
doubao-seed-2-0-pro 2026.02 6.0 81.388.870.976.679.4 45.335.439.656.644.248.8
gemini-3-flash-preview 2025.12 8.7 84.995.172.779.183.0 22.437.622.950.433.448.1
doubao-seed-2-0-lite 2026.02 5.6 78.187.670.174.777.6 37.334.330.548.837.744.2
gpt-5 2025.08 8.4 79.192.968.873.378.6 30.634.317.236.229.637.6
gpt-5-mini 2025.08 7.7 71.687.562.764.671.6 28.426.49.823.622.027.4
claude-haiku-4-5 2025.10 4.7 58.174.552.956.260.4 12.519.74.08.611.214.7
Open-Source Models
kimi-k2.5 2026.01 2.1 75.086.065.868.073.7 29.830.920.039.230.038.1
glm-4.6v 2025.12 6.2 66.777.057.958.665.1 33.221.310.426.622.930.5
qwen3-vl-235b-a22b 2025.09 2.2 66.481.563.262.568.4 23.524.213.525.821.727.6
qwen3.5-397b-a17b 2026.02 9.2 70.585.465.865.871.9 25.323.613.824.821.927.1
qwen3-vl-30b-a3b 2025.10 2.2 59.270.155.954.659.9 19.924.25.314.616.022.5
glm-4.6v-flashx 2025.12 5.8 59.371.355.455.160.3 21.621.35.712.715.319.9
qwen3.5-35b-a3b 2026.02 4.2 67.180.961.461.767.8 11.318.03.912.611.417.6

Hard = performance-conditioned AgriTaxon-Hard accuracy (≤2 of 14 models correct under the default MC protocol; 1,052 samples). The four OE track columns report exact match after normalization. EM and alias-aware Acc are unweighted macro-averages across the four tracks; the LLM judge agrees with a domain expert on 98% of 100 sampled EM errors.

πŸ“¬ Submit to Leaderboard — If you would like your model or method to appear on this leaderboard, please contact us at zengx@nercita.org.cn with your evaluation results.

Multiple Field Views Improve Recognition and Naming

AgriTaxon-Wild evaluates the same 2,969 P3151-aligned entities using one curated Commons image or one, three, and five iNaturalist field observations. Candidate names and judging procedures remain fixed, allowing a controlled comparison of image provenance and multi-view evidence.

ModelSettingMCOE-EMOE-Acc
Doubao-Seed-2.0-LiteCurated78.045.050.2
Wild K=175.636.239.9
Wild K=382.151.054.3
Wild K=584.355.759.4
Qwen3.5-35B-A3BCurated66.212.719.4
Wild K=163.213.617.6
Wild K=368.519.522.0
Wild K=570.421.323.6

A single wild view is harder than the curated image for both models. By K=3, both MC and OE-Acc exceed the corresponding curated baselines, but the seeing-without-naming gap remains at K=5.

Curated versus in-the-wild accuracy as the number of field views increases
Curated versus in-the-wild accuracy on the same 2,969 entities. Multiple field views overtake the curated baseline by K=3 for both models and both protocols.

Key Findings

Explore the Embedding Space

7,432 agricultural entities embedded by Qwen3-Embedding-0.6B and projected via t-SNE. Scroll to zoom, drag to pan, hover for details, click to open Wikipedia.

Open Full-Screen Browser β†’

Error Type Examples

Representative examples from the human-annotated 75 Acc-error cases (Gemini 3 Pro Preview, open-ended evaluation).

Taxonomic confusion example
Taxonomic Confusion (63%)
Ground truth: Stylosanthes capitata [Crop] → Prediction: Stylosanthes humilis
Same genus within Fabaceae. The model recognizes the correct genus from the pod morphology but confuses species-level features (capitulum density, beak shape).
Unrelated prediction example
Unrelated Prediction (13%)
Ground truth: Gleditsia triacanthos [Crop] → Prediction: Black walnut
Cross-family error: Fabaceae vs. Juglandaceae. Both are deciduous North American trees, but they belong to different families and are morphologically distinct.
Granularity mismatch example
Granularity Mismatch (7%)
Ground truth: Swiss Warmblood [Livestock] → Prediction: Horse
The model correctly identifies the animal as a horse but fails to specify the breed, producing a species-level answer where a breed-level answer is required.
Other Errors (17%)
No Answer (9%) — the model fails to produce any species name, typically for obscure taxa absent from training data.
Parsing Error (8%) — the model's response is truncated or malformed, outputting non-species text instead of a valid name.

How AgriTaxon Works

1

Authority-grounded data collection. We query Wikidata for agricultural entities that carry both an authoritative database identifier (FAO Ecocrop, FAO DAD-IS, or EPPO) and a Wikimedia Commons image, forming a traceable authority chain for organisms and livestock breeds.

2

Cross-domain coverage. The 48,950 queried EPPO entries—mixing pests, weeds, pathogens, host plants, and other entities—are classified by GPT-5 Mini under predefined criteria; the pest and weed assignments form the corresponding tracks. Resolution filtering removes 154 images whose shorter edge is below 224px. A livestock-specific content check removes six images that do not depict a valid animal instance.

3

Dual evaluation protocols. Multiple-choice uses text-semantically similar entity names as distractors; open-ended requires independent entity-name production. Open-ended outputs are scored by exact match and alias-aware Acc, with core summary scores macro-averaged across the four tracks.

Potential Applications

AgriTaxon is designed to support a broad range of research directions across the multimodal AI and agricultural informatics communities.

Open-Ended Visual Recognition

A testbed for models that must produce organism or breed names rather than selecting from a fixed label set.

Long-Tail Understanding

Wikipedia pageviews support analysis of the association between web visibility and recognition of rare agricultural entities.

Retrieval-Augmented Generation

Authority-grounded labels (FAO, EPPO, Wikidata QIDs) provide natural retrieval anchors for augmenting LMMs.

Agentic Reasoning

Cropping and web retrieval target complementary visual-detail and external-knowledge bottlenecks.

Agricultural AI Deployment

Pest surveillance, quarantine enforcement, crop variety verification, and livestock breed identification.

Fine-Grained Classification

Semantically hard negatives and cross-domain coverage make a challenging FGVC benchmark.

Licensing & Access

AgriTaxon is publicly accessible. Reuse of images remains subject to each image's original license and attribution requirements.

The benchmark suite is hosted on Hugging Face at Xin1818/AgriTaxon. Users should consult the per-image license and attribution metadata distributed with the dataset before reuse.

Getting Started

# Download the dataset from Hugging Face
pip install huggingface_hub
huggingface-cli download Xin1818/AgriTaxon --repo-type dataset --local-dir dataset
Loading HD from AgriTaxon HuggingFace …