Central engineering reference and operations manual for the MRDT Autonomy Software.
View the Project on GitHub MissouriMRDT/Autonomy_Software
Return to RoveSoDocs Guides for Today, Tomorrow, and Forever.
The Perception subsystem processes optical and depth imagery to identify mission targets (ArUco fiducial tags, competition props) and map physical obstacles across the terrain.
The perception pipeline integrates deep learning models, classical computer vision, stereoscopic depth mapping, and geometric projection:
[Raw Camera Stream] (ZED 2i / SIM WebRTC: 720p/1080p)
|
+-----------------------------------+
| |
v v
[TagDetector Pipeline] [ObjectDetector Pipeline]
- OpenCV ArUco (DICT_4X4_50) - LibTorch YOLO (.pt on CUDA)
- LibTorch YOLO Tag Fallback - Non-Maximum Suppression (NMS)
- CSRT / KCF Bounding Box Track - CSRT / KCF Bounding Box Track
| |
+-----------------+-----------------+
|
v
[Target Identification]
- Lifetime Thresholding (BBOX_MIN_LIFETIME_THRESHOLD)
- Screen Area Filtering (BBOX_MIN_SCREEN_PERCENTAGE)
- Target ID Matching (IdentifyTargetMarker / IdentifyTargetObject)
|
v
[Geolocation Engine]
- GeolocateBox() (src/util/vision/Geolocate.hpp)
- 3D Point Cloud Neighborhood Sampling
- Statistical 20th Percentile Depth Isolation
- Monocular Ground Plane Raycast Fallback
- UTM Coordinate Transformation via Rover Pose
|
v
[Global Waypoint Output] (geoops::Waypoint)
TagDetector.cpp)Fiducial marker detection operates through a dual-path pipeline:
cv::aruco::detectMarkers with dictionary DICT_4X4_50.cv::aruco::CORNER_REFINE_SUBPIX or CORNER_REFINE_CONTOUR) up to TAGDETECT_MAINCAM_CORNER_REFINE_MAX_ITER.TAGDETECT_MAINCAM_DETECT_INVERTED_MARKER for handling low-contrast lighting conditions.TAGDETECT_TORCH_MODEL) trained on marker silhouettes runs asynchronously via LibTorch on the GPU, outputting a candidate bounding box and confidence score (TAGDETECT_MAINCAM_TORCH_CONFIDENCE).src/util/vision/BoundingBoxTracking.cpp).constants::BBOX_MIN_LIFETIME_THRESHOLD (e.g., 0.5 seconds) and occupy at least constants::BBOX_MIN_SCREEN_PERCENTAGE of the camera frame to be considered valid targets.ObjectDetector.cpp)Target props (mallet, rock pick, water bottle) lack distinct geometric fiducials and are detected using deep neural networks:
OBJECTDETECT_TORCH_MODEL) are loaded as TorchScript (.pt) files via yolomodel::pytorch::PyTorchInterpreter.OBJECTDETECT_MAINCAM_TORCH_CONFIDENCE (e.g., 0.60) and merged using cv::dnn::NMSBoxes with an IoU threshold (OBJECTDETECT_MAINCAM_TORCH_NMS_THRESH, e.g., 0.45).tracking::MultiTracker, maintaining bounding box state during turns or temporary frame drops.Geolocate.hpp)Converting 2D pixel coordinates $(u, v)$ into 3D global UTM waypoints is performed by geoloc::GeolocateBox():
cvPointcloud (CV_32FC4).cv::Mat or cv::cuda::GpuMat) from ZED 2i or WebRTC simulation.cv::Mat formatted as CV_32FC4).geoops::RoverPose) including Easting, Northing, Altitude, and fused compass heading.geoops::Waypoint: Complete global target coordinate with UTM position, target classification type (WaypointType), and clearance radius.