a benchmark for the conditions customers actually create.
I built automated validation tooling to compare model performance across customer deployments, then connected the evaluation to a hyper-realistic synthetic orchard dataset with roughly 12 million fruit instances and expert review from six PhD agronomists.
The important part was not one score. It was building a system that could find where performance changed, explain why, and be run again.