The short answer
A plant identification app can be a useful starting point, but its result is not proof. Published evaluations show substantial variation between apps, datasets and scoring methods. An app may be right about the genus and wrong about the species, or include the correct plant among alternatives without ranking it first.
Use the result to form a candidate, then verify several independent traits. For eating, toxicity, medical, veterinary, invasive-species or pesticide decisions, get qualified local confirmation.
Why accuracy numbers disagree
“Correct” can mean different taxonomic levels
A family-level result is broader than a genus, and a genus is broader than a species. That difference matters with lookalikes. A study can report a high rate at genus level while species-level performance is much lower. Any accuracy claim should name the taxonomic level it measured.
Top result and candidate-list accuracy are different
If the correct plant appears fifth, the app has helped narrow the search but has not provided the right answer as its first choice. The 2023 PLOS ONE study used repeatable scoring systems and found considerable variation within and between apps; it also explains why rankings change when studies score results differently. Source: PLOS ONE study
The test plants shape the score
Common flowering plants in a region may be well represented in training or reference images. Rare plants, cultivars, hybrids, seedlings and plants outside the app’s strongest geography can be harder. A score from northeastern trees cannot be applied unchanged to tropical houseplants, western wildflowers or every garden worldwide.
The plant part in the photo matters
Leaves, bark, flowers, fruit and whole-plant habit provide different evidence. Rutgers researchers testing northeastern trees found that performance differed between leaf and bark photos and between leaf forms. A single number hides those differences. Source: Rutgers Urban Forestry study
Read an app result at the right level
| Result state | What it supports | What to do next |
|---|---|---|
| One candidate matches several visible traits | A working identification | Verify leaf arrangement, stem, growth habit, flower/fruit and local range |
| Several candidates share one genus | Genus may be more defensible than species | Wait for a distinguishing feature or use a botanical key/local expert |
| Candidates span unrelated groups | The image or scene is not discriminating enough | Retake with one plant, whole habit and an attached detail |
| High confidence but traits conflict | The display is not evidence of correctness | Trust the observed conflict and verify independently |
| Low confidence with one plausible suggestion | A search lead, not a confirmed name | Use it to compare references; do not save or act as if species is settled |
| Result changes with each photo | The visible evidence is unstable or incomplete | Use the disagreement checklist and capture different plant organs |
Five questions to ask before trusting the name
- Does the proposed plant grow in this region and habitat? Plausibility is not proof, but an implausible range is a reason to stop.
- Do several traits agree? Compare leaf arrangement, stem, growth habit, mature size and reproductive features—not color alone.
- Did the image show those traits? An app cannot confirm a feature that is outside the frame.
- Is the app showing alternatives or uncertainty? A candidate list can be more honest and useful than a single confident label.
- What happens if the name is wrong? The higher the consequence, the stronger the required confirmation.
How to improve the evidence
Use a clear whole-plant view to establish habit, then photograph an attached detail that separates candidates. Oklahoma State University Extension recommends focused images of the whole plant and site, close-ups of useful plant parts, and an object for scale. Include location, habitat and season in your notes. Our plant identification photo checklist turns that into a capture sequence.
What Garden Sage does not claim
Garden Sage does not publish a universal accuracy percentage for itself. Such a number would need a named, versioned, representative test set; separate genus and species scoring; regional and plant-group coverage; a fixed app/model release; and documented abstentions. Without that provenance, the responsible product behavior is to label a prediction as a prediction, keep it out of confirmed identity and plant-specific care, and ask for another view only when it would help.