The short answer

A plant identification app can be a useful starting point, but its result is not proof. Published evaluations show substantial variation between apps, datasets and scoring methods. An app may be right about the genus and wrong about the species, or include the correct plant among alternatives without ranking it first.

Use the result to form a candidate, then verify several independent traits. For eating, toxicity, medical, veterinary, invasive-species or pesticide decisions, get qualified local confirmation.

Why accuracy numbers disagree

“Correct” can mean different taxonomic levels

A family-level result is broader than a genus, and a genus is broader than a species. That difference matters with lookalikes. A study can report a high rate at genus level while species-level performance is much lower. Any accuracy claim should name the taxonomic level it measured.

Top result and candidate-list accuracy are different

If the correct plant appears fifth, the app has helped narrow the search but has not provided the right answer as its first choice. The 2023 PLOS ONE study used repeatable scoring systems and found considerable variation within and between apps; it also explains why rankings change when studies score results differently. Source: PLOS ONE study

The test plants shape the score

Common flowering plants in a region may be well represented in training or reference images. Rare plants, cultivars, hybrids, seedlings and plants outside the app’s strongest geography can be harder. A score from northeastern trees cannot be applied unchanged to tropical houseplants, western wildflowers or every garden worldwide.

The plant part in the photo matters

Leaves, bark, flowers, fruit and whole-plant habit provide different evidence. Rutgers researchers testing northeastern trees found that performance differed between leaf and bark photos and between leaf forms. A single number hides those differences. Source: Rutgers Urban Forestry study

Read an app result at the right level

Result stateWhat it supportsWhat to do next
One candidate matches several visible traitsA working identificationVerify leaf arrangement, stem, growth habit, flower/fruit and local range
Several candidates share one genusGenus may be more defensible than speciesWait for a distinguishing feature or use a botanical key/local expert
Candidates span unrelated groupsThe image or scene is not discriminating enoughRetake with one plant, whole habit and an attached detail
High confidence but traits conflictThe display is not evidence of correctnessTrust the observed conflict and verify independently
Low confidence with one plausible suggestionA search lead, not a confirmed nameUse it to compare references; do not save or act as if species is settled
Result changes with each photoThe visible evidence is unstable or incompleteUse the disagreement checklist and capture different plant organs

Five questions to ask before trusting the name

  1. Does the proposed plant grow in this region and habitat? Plausibility is not proof, but an implausible range is a reason to stop.
  2. Do several traits agree? Compare leaf arrangement, stem, growth habit, mature size and reproductive features—not color alone.
  3. Did the image show those traits? An app cannot confirm a feature that is outside the frame.
  4. Is the app showing alternatives or uncertainty? A candidate list can be more honest and useful than a single confident label.
  5. What happens if the name is wrong? The higher the consequence, the stronger the required confirmation.
Confidence is not the same as accuracy: confidence describes the system’s score for this input. Accuracy describes performance against known answers across an evaluation. A high score can still be wrong.

How to improve the evidence

Use a clear whole-plant view to establish habit, then photograph an attached detail that separates candidates. Oklahoma State University Extension recommends focused images of the whole plant and site, close-ups of useful plant parts, and an object for scale. Include location, habitat and season in your notes. Our plant identification photo checklist turns that into a capture sequence.

What Garden Sage does not claim

Garden Sage does not publish a universal accuracy percentage for itself. Such a number would need a named, versioned, representative test set; separate genus and species scoring; regional and plant-group coverage; a fixed app/model release; and documented abstentions. Without that provenance, the responsible product behavior is to label a prediction as a prediction, keep it out of confirmed identity and plant-specific care, and ask for another view only when it would help.

Sources and further reading