Benchmark Suite · 7,432 Core Entities · 4 Domains

AgriTaxon

Knowledge-grounded benchmarking of open-ended agricultural entity naming with multimodal foundation models

Xin Zeng, Benfeng Xu, Qian Chen, Jialin Kuai, Wentao Zhang, Liguo Lang, Shancheng Fang, Huarui Wu

Under Review

1,971 Crops 178 Livestock 3,485 Pests 1,798 Weeds

Naming Is the Entry Point to Precision Agriculture

Precision and smart farming increasingly rely on image-based decision systems, and those decisions begin with correctly naming what appears in a field photograph. Farmers, breeders, and quarantine officers must identify a crop variety, livestock breed, pest insect, or invasive weed before selecting a control product, managing a breed record, enforcing quarantine, monitoring biodiversity, or forecasting yield. A wrong name can therefore send the downstream decision to the wrong organism.

We call this task open-ended taxonomic naming: producing an agricultural entity name without choosing from a predefined label set. The term is used operationally here to cover recognized organism names and, for Livestock, recognized breed names. Unlike closed-set classification, it requires both visual discrimination and access to scientific names, common names, aliases, and biological relationships.

AgriTaxon is a knowledge-grounded and AI-ready benchmark suite spanning Crops, Livestock, Pests, and Weeds. Each core entity is linked through Wikidata to FAO Ecocrop, FAO DAD-IS, or the EPPO Global Database and paired with machine-readable identifiers, a canonical name, collected aliases, a domain label, and an executable evaluation protocol. We evaluate 14 multimodal foundation models using four-option recognition and open-ended naming; the alias-aware LLM judge agrees with a domain expert on 98% of 100 audited exact-match errors.

7,432
Core Entities
1,052
Hard Cases
14,845
Wild Images
14
Foundation Models

Core, Hard, and In-the-Wild Evaluation

The core benchmark measures broad agricultural entity naming. AgriTaxon-Hard is a performance-conditioned diagnostic subset, while AgriTaxon-Wild evaluates the same naming capability using multiple independent field observations.

DatasetPurposeEntitiesImages
AgriTaxon CoreBroad taxonomic-naming benchmark7,4327,432
AgriTaxon-HardPerformance-conditioned diagnostic stress test1,0521,052
AgriTaxon-WildField-condition evaluation2,96914,845

AgriTaxon-Hard contains core samples answered correctly by at most two of the 14 evaluated models under the default multiple-choice protocol. It is intended for failure analysis and tool-assisted recovery, not as an estimate of average benchmark performance.

AgriTaxon-Wild uses strict Wikidata P3151 alignment to iNaturalist, without fuzzy name matching. It retains 2,969 entities with five qualifying field photographs each and evaluates one, three, or five views of the same entity.

Leaderboard

Accuracy (%) on AgriTaxon for 14 LMMs, ranked by OE-Acc within each block. MC Macro, OE-EM, and OE-Acc are unweighted macro-averages across the four tracks; the four track-specific open-ended columns report EM. Bold = best per column.

Model Release Hard Multi-Choice Open-Ended
Crop Live. Pest Weed Macro Crop Live. Pest Weed EM Acc
Proprietary Models
gemini-3-pro-preview 2025.11 9.0 83.795.572.578.182.5 43.649.430.053.044.051.2
doubao-seed-2-0-pro 2026.02 6.0 81.388.870.976.679.4 45.335.439.656.644.248.8
gemini-3-flash-preview 2025.12 8.7 84.995.172.779.183.0 22.437.622.950.433.448.1
doubao-seed-2-0-lite 2026.02 5.6 78.187.670.174.777.6 37.334.330.548.837.744.2
gpt-5 2025.08 8.4 79.192.968.873.378.6 30.634.317.236.229.637.6
gpt-5-mini 2025.08 7.7 71.687.562.764.671.6 28.426.49.823.622.027.4
claude-haiku-4-5 2025.10 4.7 58.174.552.956.260.4 12.519.74.08.611.214.7
Open-Source Models
kimi-k2.5 2026.01 2.1 75.086.065.868.073.7 29.830.920.039.230.038.1
glm-4.6v 2025.12 6.2 66.777.057.958.665.1 33.221.310.426.622.930.5
qwen3-vl-235b-a22b 2025.09 2.2 66.481.563.262.568.4 23.524.213.525.821.727.6
qwen3.5-397b-a17b 2026.02 9.2 70.585.465.865.871.9 25.323.613.824.821.927.1
qwen3-vl-30b-a3b 2025.10 2.2 59.270.155.954.659.9 19.924.25.314.616.022.5
glm-4.6v-flashx 2025.12 5.8 59.371.355.455.160.3 21.621.35.712.715.319.9
qwen3.5-35b-a3b 2026.02 4.2 67.180.961.461.767.8 11.318.03.912.611.417.6

Hard = performance-conditioned AgriTaxon-Hard accuracy (≤2 of 14 models correct under the default MC protocol; 1,052 samples). The four OE track columns report exact match after normalization. EM and alias-aware Acc are unweighted macro-averages across the four tracks; the LLM judge agrees with a domain expert on 98% of 100 sampled EM errors.

πŸ“¬ Submit to Leaderboard — If you would like your model or method to appear on this leaderboard, please contact us at zengx@nercita.org.cn with your evaluation results.

Multiple Field Views Strengthen Recognition in Field Conditions

Field photographs vary in viewpoint, background, illumination, life stage, and image quality. AgriTaxon-Wild evaluates the same 2,969 P3151-aligned entities using one curated Commons image or one, three, and five iNaturalist field observations. Candidate names and judging procedures remain fixed, allowing a controlled test of whether corroborating field views improve recognition and naming.

ModelSettingMCOE-EMOE-Acc
Doubao-Seed-2.0-LiteCurated78.045.050.2
Wild K=175.636.239.9
Wild K=382.151.054.3
Wild K=584.355.759.4
Qwen3.5-35B-A3BCurated66.212.719.4
Wild K=163.213.617.6
Wild K=368.519.522.0
Wild K=570.421.323.6

A single wild view is harder than the curated image for both models. By K=3, both MC and OE-Acc exceed the corresponding curated baselines, but the seeing-without-naming gap remains at K=5.

Curated versus in-the-wild accuracy as the number of field views increases
Curated versus in-the-wild accuracy on the same 2,969 entities. Multiple field views overtake the curated baseline by K=3 for both models and both protocols.

Key Findings

Explore the Embedding Space

7,432 agricultural entities embedded by Qwen3-Embedding-0.6B and projected via t-SNE. Scroll to zoom, drag to pan, hover for details, click to open Wikipedia.

Open Full-Screen Browser β†’

Error Type Examples

Representative examples from the human-annotated 75 Acc-error cases (Gemini 3 Pro Preview, open-ended evaluation).

Taxonomic confusion example
Taxonomic Confusion (63%)
Ground truth: Stylosanthes capitata [Crop] → Prediction: Stylosanthes humilis
Same genus within Fabaceae. The model recognizes the correct genus from the pod morphology but confuses species-level features (capitulum density, beak shape).
Unrelated prediction example
Unrelated Prediction (13%)
Ground truth: Gleditsia triacanthos [Crop] → Prediction: Black walnut
Cross-family error: Fabaceae vs. Juglandaceae. Both are deciduous North American trees, but they belong to different families and are morphologically distinct.
Granularity mismatch example
Granularity Mismatch (7%)
Ground truth: Swiss Warmblood [Livestock] → Prediction: Horse
The model correctly identifies the animal as a horse but fails to specify the breed, producing a species-level answer where a breed-level answer is required.
Other Errors (17%)
No Answer (9%) — the model fails to produce any species name, typically for obscure taxa absent from training data.
Parsing Error (8%) — the model's response is truncated or malformed, outputting non-species text instead of a valid name.

How AgriTaxon Works

1

Authority-grounded data collection. We query Wikidata for agricultural entities that carry both an authoritative database identifier (FAO Ecocrop, FAO DAD-IS, or EPPO) and a Wikimedia Commons image, forming a traceable authority chain for organisms and livestock breeds.

2

Cross-domain coverage. The 48,950 queried EPPO entries—mixing pests, weeds, pathogens, host plants, and other entities—are classified by GPT-5 Mini under predefined criteria; the pest and weed assignments form the corresponding tracks. Resolution filtering removes 154 images whose shorter edge is below 224px. A livestock-specific content check removes six images that do not depict a valid animal instance.

3

Dual evaluation protocols. Multiple-choice uses text-semantically similar entity names as distractors; open-ended requires independent entity-name production. Open-ended outputs are scored by exact match and alias-aware Acc, with core summary scores macro-averaged across the four tracks.

Implications for Precision and Smart Farming

AgriTaxon studies the recognition step that precedes downstream agricultural decisions and highlights where practical systems need additional visual evidence or external knowledge.

Agricultural Decision Support

Evaluate whether a system can name the organism or breed before recommending control, management, or monitoring actions.

Locally Important Entities

Study why entities with low global web visibility may remain unreliable even when they matter greatly in a local production system.

Structured Knowledge Retrieval

Use FAO, EPPO, and Wikidata identifiers as traceable anchors for retrieving names, aliases, and biological relationships.

Tool-Augmented Recognition

Combine cropping for diagnostic visual detail with web retrieval when parametric knowledge is insufficient.

Field and Quarantine Workflows

Support pest surveillance, quarantine inspection, crop verification, biodiversity monitoring, and livestock breed management.

Multi-View Field Recognition

Test how several independent observations reduce ambiguity caused by viewpoint, occlusion, lighting, and life stage.

Licensing & Access

AgriTaxon is publicly accessible. Reuse of images remains subject to each image's original license and attribution requirements.

The benchmark suite is hosted on Hugging Face at Xin1818/AgriTaxon. Users should consult the per-image license and attribution metadata distributed with the dataset before reuse.

Getting Started

# Download the dataset from Hugging Face
pip install huggingface_hub
huggingface-cli download Xin1818/AgriTaxon --repo-type dataset --local-dir dataset
Loading HD from AgriTaxon HuggingFace …