Business

AV Training Data Collection Strategy: How Fleet Design and Coverage Planning Determine Model Capability

AV Training Data Collection Strategy: How Fleet Design and Coverage Planning Determine Model Capability

The quality of AV perception model training data is determined largely before any annotation begins  in the data collection strategy that decides which sensors to use, where to drive, when to drive, and how to ensure that the collected data covers the scenarios the model needs to learn from.

Most AV training data discussions focus on annotation techniques. The data collection decisions that determine what can be annotated receive less systematic attention, even though a data collection program with inadequate scenario coverage cannot produce good training data regardless of annotation quality. Annotation can only label what the collection program captured.

Sensor Suite Design: The Collection Foundation

The sensor suite on an instrumented data collection vehicle determines what physical modalities are captured and therefore what annotation is possible. The core automotive perception sensor types each contribute different information:

Cameras: Multiple cameras  typically forward-facing with varying focal lengths for near-field and far-field coverage, plus side-facing and rear-facing cameras  capture rich visual detail across the full surroundings. Camera configurations for data collection need to match the camera types and placements that the production vehicle will use, because models trained on data from a different camera configuration than the deployment platform will encounter a distribution shift at inference.

LiDAR: One or more LiDAR units covering the vehicle’s surroundings in three dimensions. Sensor placement height, rotation rate, and beam configuration all affect the point cloud characteristics, and again must match the production vehicle configuration. For data collection programs that serve multiple OEM customers with different sensor configurations, the sensor setup needs to be documented precisely so that annotation teams know the geometric parameters of the sensor setup they are annotating.

Radar: Long-range radar for highway following-distance measurement, short-range radar for low-speed proximity detection. Radar captures are particularly valuable for adverse weather scenarios where camera and LiDAR performance degrades, because the radar signal is largely unaffected by precipitation.

Reference sensors: In addition to the perception sensors, instrumented collection vehicles typically carry reference sensors whose sole purpose is providing ground truth for training data validation  high-precision GPS/IMU for precise vehicle positioning, a survey-grade LiDAR for comparison with the production LiDAR, and sometimes reference cameras with calibrated characteristics.

Geographic Coverage Planning

The geographic distribution of data collection determines the environmental and infrastructure variation the training data covers. Models trained predominantly on data from one geographic region develop perception capabilities calibrated to that region’s specific visual environment, its road marking conventions, its infrastructure styles, its vegetation types, and its typical weather conditions.

Geographic coverage planning starts with the operational design domain (ODD) the AV system is being built for. An urban robotaxi service ODD that covers specific city geographies needs data collection concentrated in those specific cities, covering the specific intersection types, road surface qualities, signage conventions, and traffic behavior patterns of those markets. A highway driving assist system ODD that covers multiple countries needs data from the specific highway corridor types of each target market.

Within the target ODD, geographic coverage planning addresses:

  • Infrastructure variety: Different intersection geometries (four-way, T-junction, roundabout, uncontrolled), different road surface types (asphalt, concrete, cobblestone), different lane marking conventions, different barrier and guardrail types, different traffic sign designs.
  • Neighborhood type variety: Urban core (high pedestrian density, complex intersection geometry, heavy cyclist presence), suburban (lower density, wider lanes, more vehicle-dominant traffic), residential (shared spaces, children present, lower speeds), commercial/industrial (heavy vehicle presence, loading zones, unusual road users).
  • Urban vs. rural: Rural routes introduce different challenges than urban ones, animals on roads, agricultural equipment, narrow lanes without markings, highway on/off-ramps.

Temporal and Condition Coverage

The time of data collection determines the environmental conditions captured. Systematic condition coverage planning ensures that the training data includes sufficient representation of conditions that are safety-critical but underrepresented in opportunistic collection.

Lighting conditions: Data collection sessions scheduled explicitly for dawn, dusk, and nighttime, in addition to daylight. The distribution of training data across lighting conditions should reflect the distribution of conditions in which the system will operate  if the system will operate 24 hours, the training data should include appropriate proportions of nighttime and transitional lighting data.

Weather conditions: Targeted collection in rain (light rain, moderate rain, heavy rain), fog, snow, and direct sun glare conditions. Weather conditions cannot be reliably scheduled in advance, requiring either large fleet operations that statistically produce adverse weather encounters or dedicated collection campaigns that follow weather events to specific locations.

Seasonal variation: Road surface conditions, vegetation density, sun angle, and in some markets the presence of snow cover on road markings all vary seasonally. Training data collected across all seasons is more representative of the full operational range than data collected in a single season.

Traffic conditions: Peak-hour congestion, off-peak light traffic, construction-reduced capacity, event-related congestion, and overnight sparse traffic each present different detection, prediction, and planning challenges. Condition-targeted collection sessions during specific traffic states ensure representation across the traffic density range.

Rare Event and Edge Case Collection

The events that determine AV system safety are disproportionately the events that are rare in normal collection, the scenarios that occur at frequencies of once per several thousand hours of driving but that carry high safety consequences when the system encounters them.

Targeted collection strategies for rare events:

Geographic targeting: Some rare road configurations appear frequently in specific locations  unusual intersection geometries, complex interchange structures, grade crossings, drawbridge approaches. Targeted collection at locations where specific rare configurations appear efficiently accumulates examples that would be under-collected by random geographic coverage.

Temporal targeting: Some rare events are predictable by time  construction zones appear during specific working hours, school zones have elevated pedestrian activity during school start and end times, event venues generate unusual traffic and pedestrian patterns during events. Time-targeted collection during these predictable rare event windows efficiently accumulates the scenarios that random temporal coverage would rarely encounter.

Scenario induction in closed environments: For rare events too dangerous to encounter on public roads during data collection  wrong-way vehicles, sudden road surface hazards, emergency vehicle scenarios  closed-environment collection with scenario actors recreates the event in a controlled setting where it can be captured safely.

Operational data mining: Vehicles in early deployment generate operational data that, at fleet scale, encounters rare events at frequencies proportional to operational hours times the natural event rate. Mining this operational data for the rare events that standard collection wouldn’t produce efficiently and combining those examples with standard collection data ensures rare event coverage scales with deployment experience.

Data Volume vs. Data Diversity: The Critical Tradeoff

AV programs with limited collection budgets face a fundamental tradeoff: should the collection program maximize volume (more hours in the same environments) or diversity (fewer hours across more environments and conditions)?

The volume-maximizing choice produces a large dataset of similar examples. The diversity-maximizing choice produces a smaller dataset of varied examples. Research and production experience consistently support the diversity-maximizing approach for AV Perception Model Training Data model quality: a model trained on 100 hours of diverse data outperforms a model trained on 1,000 hours of similar data on out-of-distribution scenarios  which are exactly the scenarios that determine real-world safety.

This doesn’t mean volume doesn’t matter. It means that collection effort invested in diversity produces more model improvement per hour than collection effort invested in additional repetition of already-covered scenarios. The optimal collection strategy covers the target diversity efficiently and then scales volume within the covered diversity, rather than scaling volume within a narrow slice of the coverage space.

Final Thought

AV perception model training data quality begins with collection strategy. The sensor suite design, geographic coverage plan, condition coverage plan, and rare event collection strategy together determine what the annotation program has to work with  and no annotation quality, however high, can produce training data for scenarios the collection program never captured.

Programs that design collection strategy with the same rigor they apply to annotation programs build training datasets whose coverage matches the operational reality of the AV system they support. Programs that collect opportunistically and annotate thoroughly produce well-labeled examples of common scenarios and undertrained models on the rare ones.

About Author

tienblackwood

Leave a Reply

Your email address will not be published. Required fields are marked *