Knowledge-grounded benchmarking of open-ended agricultural entity naming with multimodal foundation models
Under Review
Precision and smart farming increasingly rely on image-based decision systems, and those decisions begin with correctly naming what appears in a field photograph. Farmers, breeders, and quarantine officers must identify a crop variety, livestock breed, pest insect, or invasive weed before selecting a control product, managing a breed record, enforcing quarantine, monitoring biodiversity, or forecasting yield. A wrong name can therefore send the downstream decision to the wrong organism.
We call this task open-ended taxonomic naming: producing an agricultural entity name without choosing from a predefined label set. The term is used operationally here to cover recognized organism names and, for Livestock, recognized breed names. Unlike closed-set classification, it requires both visual discrimination and access to scientific names, common names, aliases, and biological relationships.
AgriTaxon is a knowledge-grounded and AI-ready benchmark suite spanning Crops, Livestock, Pests, and Weeds. Each core entity is linked through Wikidata to FAO Ecocrop, FAO DAD-IS, or the EPPO Global Database and paired with machine-readable identifiers, a canonical name, collected aliases, a domain label, and an executable evaluation protocol. We evaluate 14 multimodal foundation models using four-option recognition and open-ended naming; the alias-aware LLM judge agrees with a domain expert on 98% of 100 audited exact-match errors.
The core benchmark measures broad agricultural entity naming. AgriTaxon-Hard is a performance-conditioned diagnostic subset, while AgriTaxon-Wild evaluates the same naming capability using multiple independent field observations.
| Dataset | Purpose | Entities | Images |
|---|---|---|---|
| AgriTaxon Core | Broad taxonomic-naming benchmark | 7,432 | 7,432 |
| AgriTaxon-Hard | Performance-conditioned diagnostic stress test | 1,052 | 1,052 |
| AgriTaxon-Wild | Field-condition evaluation | 2,969 | 14,845 |
AgriTaxon-Hard contains core samples answered correctly by at most two of the 14 evaluated models under the default multiple-choice protocol. It is intended for failure analysis and tool-assisted recovery, not as an estimate of average benchmark performance.
AgriTaxon-Wild uses strict Wikidata P3151 alignment to iNaturalist, without fuzzy name matching. It retains 2,969 entities with five qualifying field photographs each and evaluates one, three, or five views of the same entity.
Accuracy (%) on AgriTaxon for 14 LMMs, ranked by OE-Acc within each block. MC Macro, OE-EM, and OE-Acc are unweighted macro-averages across the four tracks; the four track-specific open-ended columns report EM. Bold = best per column.
| Model | Release | Hard | Multi-Choice | Open-Ended | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Crop | Live. | Pest | Weed | Macro | Crop | Live. | Pest | Weed | EM | Acc | |||
| Proprietary Models | |||||||||||||
gemini-3-pro-preview |
2025.11 | 9.0 | 83.7 | 95.5 | 72.5 | 78.1 | 82.5 | 43.6 | 49.4 | 30.0 | 53.0 | 44.0 | 51.2 |
doubao-seed-2-0-pro |
2026.02 | 6.0 | 81.3 | 88.8 | 70.9 | 76.6 | 79.4 | 45.3 | 35.4 | 39.6 | 56.6 | 44.2 | 48.8 |
gemini-3-flash-preview |
2025.12 | 8.7 | 84.9 | 95.1 | 72.7 | 79.1 | 83.0 | 22.4 | 37.6 | 22.9 | 50.4 | 33.4 | 48.1 |
doubao-seed-2-0-lite |
2026.02 | 5.6 | 78.1 | 87.6 | 70.1 | 74.7 | 77.6 | 37.3 | 34.3 | 30.5 | 48.8 | 37.7 | 44.2 |
gpt-5 |
2025.08 | 8.4 | 79.1 | 92.9 | 68.8 | 73.3 | 78.6 | 30.6 | 34.3 | 17.2 | 36.2 | 29.6 | 37.6 |
gpt-5-mini |
2025.08 | 7.7 | 71.6 | 87.5 | 62.7 | 64.6 | 71.6 | 28.4 | 26.4 | 9.8 | 23.6 | 22.0 | 27.4 |
claude-haiku-4-5 |
2025.10 | 4.7 | 58.1 | 74.5 | 52.9 | 56.2 | 60.4 | 12.5 | 19.7 | 4.0 | 8.6 | 11.2 | 14.7 |
| Open-Source Models | |||||||||||||
kimi-k2.5 |
2026.01 | 2.1 | 75.0 | 86.0 | 65.8 | 68.0 | 73.7 | 29.8 | 30.9 | 20.0 | 39.2 | 30.0 | 38.1 |
glm-4.6v |
2025.12 | 6.2 | 66.7 | 77.0 | 57.9 | 58.6 | 65.1 | 33.2 | 21.3 | 10.4 | 26.6 | 22.9 | 30.5 |
qwen3-vl-235b-a22b |
2025.09 | 2.2 | 66.4 | 81.5 | 63.2 | 62.5 | 68.4 | 23.5 | 24.2 | 13.5 | 25.8 | 21.7 | 27.6 |
qwen3.5-397b-a17b |
2026.02 | 9.2 | 70.5 | 85.4 | 65.8 | 65.8 | 71.9 | 25.3 | 23.6 | 13.8 | 24.8 | 21.9 | 27.1 |
qwen3-vl-30b-a3b |
2025.10 | 2.2 | 59.2 | 70.1 | 55.9 | 54.6 | 59.9 | 19.9 | 24.2 | 5.3 | 14.6 | 16.0 | 22.5 |
glm-4.6v-flashx |
2025.12 | 5.8 | 59.3 | 71.3 | 55.4 | 55.1 | 60.3 | 21.6 | 21.3 | 5.7 | 12.7 | 15.3 | 19.9 |
qwen3.5-35b-a3b |
2026.02 | 4.2 | 67.1 | 80.9 | 61.4 | 61.7 | 67.8 | 11.3 | 18.0 | 3.9 | 12.6 | 11.4 | 17.6 |
Hard = performance-conditioned AgriTaxon-Hard accuracy (≤2 of 14 models correct under the default MC protocol; 1,052 samples). The four OE track columns report exact match after normalization. EM and alias-aware Acc are unweighted macro-averages across the four tracks; the LLM judge agrees with a domain expert on 98% of 100 sampled EM errors.
π¬ Submit to Leaderboard — If you would like your model or method to appear on this leaderboard, please contact us at zengx@nercita.org.cn with your evaluation results.
Field photographs vary in viewpoint, background, illumination, life stage, and image quality. AgriTaxon-Wild evaluates the same 2,969 P3151-aligned entities using one curated Commons image or one, three, and five iNaturalist field observations. Candidate names and judging procedures remain fixed, allowing a controlled test of whether corroborating field views improve recognition and naming.
| Model | Setting | MC | OE-EM | OE-Acc |
|---|---|---|---|---|
| Doubao-Seed-2.0-Lite | Curated | 78.0 | 45.0 | 50.2 |
| Wild K=1 | 75.6 | 36.2 | 39.9 | |
| Wild K=3 | 82.1 | 51.0 | 54.3 | |
| Wild K=5 | 84.3 | 55.7 | 59.4 | |
| Qwen3.5-35B-A3B | Curated | 66.2 | 12.7 | 19.4 |
| Wild K=1 | 63.2 | 13.6 | 17.6 | |
| Wild K=3 | 68.5 | 19.5 | 22.0 | |
| Wild K=5 | 70.4 | 21.3 | 23.6 |
A single wild view is harder than the curated image for both models. By K=3, both MC and OE-Acc exceed the corresponding curated baselines, but the seeing-without-naming gap remains at K=5.
7,432 agricultural entities embedded by Qwen3-Embedding-0.6B and projected via t-SNE. Scroll to zoom, drag to pan, hover for details, click to open Wikipedia.
Representative examples from the human-annotated 75 Acc-error cases (Gemini 3 Pro Preview, open-ended evaluation).
Authority-grounded data collection. We query Wikidata for agricultural entities that carry both an authoritative database identifier (FAO Ecocrop, FAO DAD-IS, or EPPO) and a Wikimedia Commons image, forming a traceable authority chain for organisms and livestock breeds.
Cross-domain coverage. The 48,950 queried EPPO entries—mixing pests, weeds, pathogens, host plants, and other entities—are classified by GPT-5 Mini under predefined criteria; the pest and weed assignments form the corresponding tracks. Resolution filtering removes 154 images whose shorter edge is below 224px. A livestock-specific content check removes six images that do not depict a valid animal instance.
Dual evaluation protocols. Multiple-choice uses text-semantically similar entity names as distractors; open-ended requires independent entity-name production. Open-ended outputs are scored by exact match and alias-aware Acc, with core summary scores macro-averaged across the four tracks.
AgriTaxon studies the recognition step that precedes downstream agricultural decisions and highlights where practical systems need additional visual evidence or external knowledge.
Evaluate whether a system can name the organism or breed before recommending control, management, or monitoring actions.
Study why entities with low global web visibility may remain unreliable even when they matter greatly in a local production system.
Use FAO, EPPO, and Wikidata identifiers as traceable anchors for retrieving names, aliases, and biological relationships.
Combine cropping for diagnostic visual detail with web retrieval when parametric knowledge is insufficient.
Support pest surveillance, quarantine inspection, crop verification, biodiversity monitoring, and livestock breed management.
Test how several independent observations reduce ambiguity caused by viewpoint, occlusion, lighting, and life stage.
AgriTaxon is publicly accessible. Reuse of images remains subject to each image's original license and attribution requirements.
The benchmark suite is hosted on Hugging Face at Xin1818/AgriTaxon. Users should consult the per-image license and attribution metadata distributed with the dataset before reuse.
# Download the dataset from Hugging Face pip install huggingface_hub huggingface-cli download Xin1818/AgriTaxon --repo-type dataset --local-dir dataset