Ezina · A Field Manual for Plant Care · Contiguous U.S. Vol. I · Ed. 2027 · Proof copy

The truth on plant ID apps

No publicly accessible research shows how effective Plant ID apps are, so we’re setting the bar to help you be informed.

You point a plant ID app at the plant at the store, or the unexpected growth in your garden. It gives you a name in seconds, often with a confidence percentage attached. But what are these numbers actually based on?

How these apps get scientifically evaluated and how someone like you or me would actually use them are two very different things. The studies that exist cluster around a few specific regions and plant groups, none of them the plants most people actually grow. For disease identification, the standard training dataset (PlantVillage) is flawed in a way that makes lab results look better than real ones1,2. And for PictureThis, Planta, Seek, and the identification engines behind Ezina (Plant.id and Pl@ntNet, with Claude Haiku as backup), no peer-reviewed study has measured how often any of them correctly names a Monstera deliciosa from the nursery, a weed in your yard, or a tomato seedling on your kitchen windowsill. That problem is exactly why this piece exists.

Why I’m publishing this

I’m building Ezina, a plant care app launching in early 2027. I’m writing this because the way accuracy gets talked about around plant apps is misleading, and I want to lay the evidence out honestly before I make any claims of my own. Transparency is one of Ezina’s founding principles, and the fairest place to prove it is a study bringing our tools against the competitors so we can all see how these tools work for us.

To accomplish that this piece does two things.

First, I walk through what’s known and what isn’t about how accurate these apps really are. The research, the marketing claims, and the space between them, all cited at the bottom so you can check my work.

Second, I commit publicly to a measurement protocol, locked in writing before I’ve run a single identification. I’m publishing the method before I have any data, so I can’t adjust either one after the fact to make Ezina look good.

What each app publicly says about accuracy

Each of the four apps handles the accuracy question differently. Here’s what each one publishes:

App What they say about accuracy Methodology shown Where you can find it
PictureThis “Over 98% accuracy” across “over 400,000 plant species” None iOS App Store[3]
Planta No number at all N/A Homepage, About, support, App Store all describe ID only as “instantly” or “in seconds”[4]
Seek (iNaturalist) Paired bar charts comparing model versions Partial: 1,000 held-out Research Grade observations per taxonomic group Model release blog posts[5]
Plant.id (Kindwise) “93% top-3 accuracy”; 35,000+ classes; regional v4.0.1 top-1 gains over the prior model Partial: 50,000 GBIF observations Product page and model update blog[6]
Ezina Nothing yet This document

The spread itself highlights the problem. PictureThis publishes the highest headline number and the least to back it up. Planta publishes no number at all. Seek and Plant.id sit in between, with partial methodology disclosed (Plant.id names a dataset and a sample size; Seek describes the held-out observation count and the protocol). Neither one publishes a quotable accuracy figure on the categories most people actually photograph.

The PictureThis “98% accuracy” figure deserves a closer look, because the app is one of the most prominent in the space and they quote the number often in their marketing materials. No test dataset is named. No app version is stamped. No scoring rule is defined (top-one, top-k, exact match, partial credit). PictureThis has been measured at 97.27% genus-level accuracy on leaf photographs in a single narrowly scoped peer-reviewed study of Northeast US trees7, close to but not at the headline 98% claim. No other public research backs the broader figure across 400,000 species, which is in large part why this piece exists. Without methodology disclosure, there is no way to validate published accuracy numbers, and consumers can be misled by them whether or not anyone intends that outcome.

In 2024, Rzanny et al. re-ran the Hart photo set, this time with the developers of one of the tested apps (Flora Incognita) re-adjudicating disputed identifications10. Flora Incognita’s accuracy rose to 98.8% in the re-evaluation. All seven authors are members of the Flora Incognita project, and the paper carries an explicit conflict-of-interest statement.

The 98.8% figure circulates widely in comparative marketing without that context. The point worth making is methodological: accuracy figures from developer-controlled re-adjudication aren’t equivalent to independent measurement, even when the venue is peer-reviewed.

What independent research has measured

A small body of peer-reviewed research has actually tested these apps. Five studies are worth knowing, because together they’re the entire independent evidence base, as of mid-2026, for the apps in this comparison.

The earliest is Jones (2020), an AoB PLANTS paper testing nine plant ID apps on 38 photographs of British flora8. Plant.id ranked first, identifying the correct species on the first attempt 57% of the time. This is still the only peer-reviewed evaluation of Plant.id, one of the engines behind Ezina. The sample is small. The flora is regional. It’s what we have.

Three years later, Hart et al. (2023) tested five free apps on 857 expert-identified photographs covering 277 species9. Top-one accuracy ranged from 86.9% for LeafSnap and 86.5% for Pl@ntNet down to 65.6% for Seek, 57.2% for Google Lens, and 46.4% for PlantSnap. PictureThis, Planta, and Plant.id weren’t included; the study only looked at free apps.

In 2022, a team at the Rutgers Urban Forestry Program tested six apps on 440 photographs of 55 tree species common to New Jersey7. PictureThis came out on top: 97.27% genus accuracy and 83.86% species accuracy on leaf photographs, with first-suggestion scoring. Bark photographs were substantially harder for every app. PictureThis dropped to 65.45% genus and 51.82% species. The pattern held across all six apps: leaves outperformed bark, and genus accuracy was substantially higher than species accuracy. This is the strongest independent evidence on PictureThis that exists, and its scope is one US state’s tree flora.

PictureThis was also tested in 2020 on 17 toxic plant species in a Journal of Medical Toxicology study11. Aggregate accuracy across 457 observations was 96%. But species-level accuracy was perfect for only 10 of the 17 species in the test set.

Kuchařová et al. (2025) most recently tested Seek on 40 conifer species at the Botanical Garden of Olomouc in the Czech Republic12. Seek hit 92.5% genus-level accuracy but only 38.75% species-level accuracy. Common European natives like Picea abies, Pinus sylvestris, Larix decidua, and Pseudotsuga menziesii were identified perfectly. Rare and non-native cultivated conifers failed badly. The pattern is consistent across the literature: common species do well; less-common cultivars don’t.

Five studies. Roughly 1,400 unique photographs covering on the order of 380 species, with the evidence base heavily concentrated on British flora, Northeast US trees, and conifers, plus a small toxic-plants subset. That represents roughly 0.1% of the global vascular plant species described in open taxonomic databases like GBIF.

What the literature doesn’t cover

Your average use case likely isn’t going to be British wild flora, Northeast US tree species, or Czech conifers. What’s in your house and yard right now, or what you encounter in your local area, isn’t represented in these studies. It’s a problem from a trust standpoint, and from a business one. If these companies aren’t sharing how well their tools perform in your day-to-day usage, it can easily come across as misleading.

The categories that drive most consumer plant ID requests are the categories with the least published evidence.

For PictureThis specifically, the two studies that tested it together cover about 72 species. PictureThis claims to identify 400,000. That’s 0.018% of the claimed universe, with a non-random sample weighted toward trees and toxic plants.

Accuracy on the remaining 399,928 species, including the houseplants and common garden cultivars people open the app to identify, is unmeasured.

The same gap applies to every app in this comparison. Plant.id was measured on 38 photographs of British plants. Pl@ntNet has more coverage, mostly through citizen-science datasets of wild flora. Seek’s published data is wild flora plus conifers. Planta has no published independent evaluation at all.

The gap is structural. It comes from how research priorities work. Independent evaluations get conducted by botanists and forestry researchers, on populations that matter to botany and forestry. The plants that drive the bulk of consumer ID requests (houseplants, weeds, garden-center cultivars) aren’t the plants the research community has been paid to test.

A short note on disease ID. The gap is even larger for plant disease identification, because the standard public training dataset (PlantVillage) has known methodological problems. Models trained on PlantVillage routinely lose 30 to 70 percentage points of accuracy when you test them on field photographs instead of the lab images PlantVillage contains1,2. Any consumer disease ID app advertising 99% accuracy is, in all likelihood, testing on lab images rather than on photos from actual users. Ezina’s disease ID runs on Plant.id, and any results I publish will be held to the same standard as everything else: independent, methodology disclosed, no headline numbers without context.

What I’m committing to

A note about Ezina’s identification first. I’m currently blending responses from Plant.id and Pl@ntNet, both licensed third-party engines. Claude Haiku, Anthropic’s multimodal model13, is being wired in as a secondary check on cases where that blend struggles. Longer-term, I’m working toward a proprietary model trained on the gap categories this piece is about: houseplants, cultivars, seedlings, weeds. The protocol below tests what’s actually shipping when I run it. Anything I might ship later won’t count toward those numbers.

I’m going to test four apps (PictureThis, Planta, Seek, and Ezina) on 75 to 100 photographs drawn from the categories the existing research doesn’t cover. The protocol is being published before any identification is run, so the methodology is fixed by what I said I’d do rather than adjusted to match what I wished I’d found.

Category Photo count Why this category Ground truth source
Houseplants 25–30 Biggest evidence gap, and the primary use case for most consumers Tagged retail nursery stock, verified personal collections, hobbyist confirmation
Cultivars 15–20 The hardest case in plant ID, where most apps fail Seed catalog reference photos with documented identity; tagged nursery cultivar stock
Seedlings 10–15 Pre-flower, pre-fruit plants are a known weak spot for image recognition models Vegetable seedling reference photos with known cultivar
Weeds 15–20 Common consumer use case, underrepresented in the research UC IPM, UGA Bugwood, Cooperative Extension photo libraries
Northeast US trees 10–15 Calibration block, so I can cross-check methodology against Schmidt et al. 2022[7] Schmidt et al. supplementary archive[7]

Each photograph is uploaded to all four apps within a single seven-day window, so the comparison isn’t muddied by an app updating its model mid-test. App build numbers are recorded on the day of testing.

The primary accuracy measurement uses photos with EXIF location data stripped, so no app gets an advantage from geographic priors. A subset of 30 photos drawn proportionally across the five categories is also tested with each photo’s actual capture location embedded as EXIF, capturing how much geographic priors actually move accuracy in each app. The EXIF effect is reported per app alongside the primary numbers.

Seek’s identification is captured through the iNaturalist app’s photo-upload feature when Seek’s consumer app does not support gallery upload directly. The underlying computer vision model is identical between Seek and iNaturalist.

Each identification is scored three ways. Correct means the genus, species, and cultivar (where applicable) all match. Partial means the genus is right but the species or cultivar is wrong. Wrong means the genus is wrong, or no identification was produced.

Two raters independently score every result. They’re blinded to which app produced which output. I’ll report inter-rater agreement as Cohen’s kappa. Where the raters disagree, the ground truth source resolves it, not me.

The headline accuracy I publish will be Correct-only percentage. Partial-credit accuracy will be in an appendix, so anyone who wants to recompute under a looser definition can.

Sample sizes are small by design. The protocol is a first measurement of categories that haven’t been measured before. I won’t frame the results as a definitive benchmark of which app is best, and I’ll report confidence intervals next to every headline figure. The EXIF-effect numbers are measured on a 30-photo subset and should be read as rough estimates rather than precise effect sizes.

Results will be posted by February 15, 2027. If Ezina performs poorly on any category, those results will be published with the same prominence and detail as anything else.

What I’m not claiming

Three things to be clear about, so this piece doesn’t get read into something it isn’t.

The methodology critique of PictureThis isn’t a charge of dishonesty. PictureThis has been part of more independent measurement than any other paid app here. The undisclosed methodology behind their “98%” headline is a real problem, but it’s a different problem from Planta publishing no quantitative claim at all, and a smaller one.

The absence of evidence on Planta isn’t a claim that Planta’s identification works poorly. It might work fine. Nobody outside the company has data on it (including me), because no independent evaluation has been published. The absence is just an absence.

I’m also not claiming Ezina’s identification is more accurate than the competition. The data to back that claim doesn’t exist yet. I won’t publish a comparative accuracy number until it does, and until I can defend the methodology behind the comparison.

What I hope you take from this

If you’re trying to decide which plant ID app to use, the honest answer is that no consumer plant ID app has been publicly measured on the photographs you’re most likely to take. The most thorough independent test of PictureThis covers 55 Northeast US tree species. Plant.id has 38 photographs of British plants. Planta has no independent test at all. Seek has wild flora and conifers.

What I’m hoping this piece supports is the case for methodology disclosure as a real criterion when you choose among consumer plant apps. It’s the one I’m asking to have Ezina judged on, and the one I’m publicly inviting you to hold me to.

Results from the protocol above will post by February 15, 2027, in the same place this piece lives. If the numbers come out poorly on any category, you’ll see them there with everything else.


References

  1. 1. Xu, M., Park, J.-E., Lee, J., Yang, J., & Yoon, S. (2024). Plant disease recognition datasets in the age of deep learning: Challenges and opportunities. Frontiers in Plant Science, 15, 1452551. https://doi.org/10.3389/fpls.2024.1452551
  2. 2. Noyan, M. A. (2022). Uncovering bias in the PlantVillage dataset (arXiv:2206.04374). arXiv. https://arxiv.org/abs/2206.04374
  3. 3. PictureThis (Glority, LLC). App Store listing, accessed 2026-06-06. https://apps.apple.com/us/app/picturethis-plant-identifier/id1252497129
  4. 4. Planta. Homepage, About, support, App Store listing, accessed 2026-06-06. https://getplanta.com/
  5. 5. iNaturalist. “Updated computer vision model and geomodel with over 1,500 new taxa,” post 118979, accessed 2026-06-06. https://www.inaturalist.org/posts/118979
  6. 6. Kindwise / Plant.id. Product page and model update posts, accessed 2026-06-06. https://www.kindwise.com/plant-id
  7. 7. Schmidt, R.J., Casario, B.M., Zipse, P.C., Grabosky, J.C. (2022). “An Analysis of the Accuracy of Photo-Based Plant Identification Applications on Fifty-Five Tree Species.” Arboriculture & Urban Forestry, 48(1), 27–43. doi:10.48044/jauf.2022.003
  8. 8. Jones, H.G. (2020). “What plant is that? Tests of automated image recognition apps for plant identification on plants from the British flora.” AoB PLANTS, 12(6), plaa052. doi:10.1093/aobpla/plaa052
  9. 9. Hart, A.G., et al. (2023). “Assessing the accuracy of free automated plant identification applications.” People and Nature. doi:10.1002/pan3.10460
  10. 10. Rzanny, M., et al. (2024). Re-evaluation of plant identification app accuracy. People and Nature. doi:10.1002/pan3.10676. All seven authors are affiliated with the Flora Incognita project; the paper includes an explicit conflict-of-interest statement.
  11. 11. Otter, J., Mayer, S., Tomaszewski, C.A. (2020). “Plant identification for emergency clinical toxicology applications.” Journal of Medical Toxicology. doi:10.1007/s13181-020-00803-6
  12. 12. Kuchařová, Z., Vašutová, D., Vašut, R.J. (2025). “Evaluating the accuracy of the Seek app for conifer identification: A baseline for future educational use.” Natural Sciences Education. doi:10.1002/nse2.70030
  13. 13. Anthropic. Claude models documentation, accessed 2026-06-06. https://docs.anthropic.com/en/docs/about-claude/models