How LightlyTrain Powers Vale's Aerial Wildlife Detection Models with Less Data
Lightly helped Vale train accurate aerial wildlife detectors with DINOv3, solving the scarce labeled data problem that field teams can't avoid.
Lightly helped Vale train accurate aerial wildlife detectors with DINOv3, solving the scarce labeled data problem that field teams can't avoid.
.png)

Lightly helped Vale train accurate aerial wildlife detectors with DINOv3, solving the scarce labeled data problem that field teams can't avoid.

Vale builds an end-to-end wildlife monitoring ecosystem: long-range wireless trail cameras, vehicle-mounted optical survey payloads, and the Vale Platform that turns raw imagery into habitat insights with population counts and movement patterns.Â
Vale works with state wildlife agencies including Colorado Parks & Wildlife, and Texas Parks & Wildlife, as well as land conservation groups like the East Foundation, on species ranging from elk and bighorn sheep to endangered ocelots and nilgai.
Millions of Images, Not Enough Labels, No Pipeline to Close to GapÂ
Vale's constraints are familiar to any team scaling computer vision in the field: image volume outpaces manual review, labeled data is scarce, and building a training pipeline that works well on limited labels is its own engineering problem. Three specific issues stood in the way:Â
1. More footage than any team can reviewÂ
Wildlife population surveys generate more imagery than any team of biologists can review by hand. Colorado's aerial surveys alone produce gigabyte-sized images that break down into roughly 1,000 sub-images each, and a biologist needs about 10 minutes to work through just one.Â
2. Not enough labeled data to train onÂ
At the same time, publicly available aerial labeled datasets for this type of work do not exist. The team needed a way to train accurate detectors on the modest, often highly specialized datasets they can realistically label themselves, for species and situations with little to no existing public training data. That constraint pointed them toward LightlyTrain’s pretraining.Â
3. Building the pipeline in-house is a heavy liftÂ
Vale had already been building on DINOv2 for keypoint and landmark estimation, and had previously built detector integrations themselves using architectures from Mask R-CNN through classic CNN-based detectors.Â
When DINOv3 was released, the team wanted it as an upgrade path for their detectors, particularly for aerial detection. But building and maintaining that training pipeline in-house was a significant lift, and it is not the work a three-person ML team wants to own.Â
One of our main constraints is that we don't have millions of labeled images, just a small amount. Using LightlyTrain to train a Vision Transformer backbone like DINOv2 or DINOv3 gets us way more efficient training results even with less data.
DINOv3 Detectors Built Directly on Vale's Existing PipelineÂ
Vale found LightlyTrain through its GitHub repository and started testing it directly, drawn by how easily it dropped into their existing Python workflow.Â
That initial test turned into Vale adopting LightlyTrain to train DINOv3-based object detectors across their aerial and trail camera use cases.Â
For their Colorado Parks & Wildlife aerial elk survey work, the team built a synthetic dataset of over 250,000 images and fine-tuned it on 2,000 real labeled samples, an approach made viable by the efficiency of the DINOv3 backbone on limited real-world data.Â
Vale also pointed to the reliability of the training process itself as a differentiator from building it in-house:Â
Lightly knows how to train these models properly and how to iterate on things. The process is so straightforward with LightlyTrain, it's refreshing. I don't have to go through and do all my own compositions, though I can still tweak it if I want. Even with an engineering background, it made this so efficient, it removed so much pain right off the bat. - Joseph Porter, Co-Founder, ValeÂ
The team is also planning to run LightlyTrain's distillation process on an upcoming multi-terabyte aerial imagery delivery.Â
Results: Outperforming MegaDetector on a Fraction of the DataÂ
In summary, with LightlyTrain, Vale has:Â
- Reached roughly 82% precision on herd size estimation for Colorado's aerial elk surveys on first large scale pass, using a DINOv3-based detector fine-tuned on 250,000 synthetic images plus 2,000 real labeled samplesÂ
- Built the foundation for automated population survey analysis that Vale estimates would otherwise take biologists roughly 16,000 hours to complete by hand on a single large aerial datasetÂ
Looking ahead, Vale is expanding into individual animal identification, including unique ID for white-tailed deer, bighorn sheep, and mountain lions in Colorado, and is preparing to launch its long-range wireless trail camera system, with edge detection as a future area of exploration with Lightly.

Self-Supervised Pretraining
Leverage self-supervised learning to pretrain models
AI Training Data for LLMs & CV
Expert training data services for LLMs, AI Agents and vision


Picking DINOv3 or YOLO11 is easy. Getting it to run in production isn’t.
Learn how to do it properly. 👇