5th Workshop on Maritime Computer Vision (MaCVi)
MaCVi @ WACV 2027
Challenges / Diver3D sonar detection
Diver3D: 3D Sonar Diver Detection
Detect divers as full 9-DoF oriented boxes from sparse 3D sonar point clouds.
An autonomous underwater vehicle assisting a human diver must track not just where the diver is, but how their body is oriented — and underwater, cameras fail first. Forward-looking sonar survives turbidity but projects the scene onto a 2D plane, discarding the elevation that orientation depends on. Recently commercialised 3D sonar keeps elevation, at the price of sparse, noisy, multipath-corrupted returns that dense-LiDAR detectors were never designed for. And unlike cars and pedestrians, a free-swimming diver pitches and rolls continuously: the yaw-only box used throughout automotive 3D detection cannot represent the pose at all.
The same diver, four ways. A yaw-only box with the correct centre and extents still cannot enclose a diver tilted 72° from vertical — that is the gap this challenge measures.
Download Diver3D v2 (3.45 GB) SHA-256 checksum
Submit predictions Leaderboard
The link opens a shared folder containing diver3d_v2.zip
and diver3d_v2.zip.sha256. Expected SHA-256 of the zip:
73bcee1251f5f12932859e4ef421e646c569124211628633989d4303aa5f89cb — verify with sha256sum -c diver3d_v2.zip.sha256.
Quick links: Task Dataset Box convention Evaluation Submission Ask for help
Task
Given one 3D sonar point cloud — an (N, 4) array of x, y, z, intensity
in the sonar frame, with N between 428 and 11,965 — output every diver as a
9-DoF oriented box: a 3D centre, extents (length, width, height),
a full rotation R ∈ SO(3), and a confidence score. There is one class
(Diver); 27.6 % of frames contain no diver and are scored as real negatives.
Frames are provided in acquisition order with stable instance ids on the training labels, so
temporal methods are allowed, but the task is scored per frame.
Dataset
Diver3D was collected at Blue Grotto, a natural spring cave-diving site in Williston, Florida,
across seven field sessions. Divers in a natural site exhibit the pitch and roll rarely seen in
pool recordings: 34.8 % of annotated instances are near-upright, 30.7 % near-horizontal,
and the rest in between. Sonar frames are annotated by hand with a full-SO(3) box
enclosing the complete perceived body extent, using synchronised camera imagery as the annotation
aid. The camera imagery is not part of the challenge release; this is a sonar-only task.
| split | scenes | frames | annotated divers | labels |
|---|---|---|---|---|
| train | 21 | 20,861 | 19,482 | released |
| val | 7 | 8,037 | 7,509 | released |
| test | 8 | 8,112 | 7,577 | withheld |
| total | 36 | 37,010 | 34,568 |
The split is scene-wise, so temporally adjacent frames never straddle a split. Four divers
appear in the data and every split contains at least three of them, so the same people are seen
in train, val and test. That is a property of the dataset: results measure detection of
diver bodies in this environment, not generalisation to unseen individuals, and should be cited
as such. Point clouds span
x ∈ [0.1, 16.0], y ∈ [−11.3, 11.3],
z ∈ [−5.5, 5.5] m; the organisers' baseline crops to
[0, 12] × [−5, 5] × [−2.5, 2.5] m, which keeps 99.9 % of
annotated divers, but any crop is allowed and evaluation is over all annotated instances.
Box convention
This is where sonar 3D detection goes wrong most often, so it is stated precisely. The sonar
frame is right-handed in metres: x forward along the boresight, y left,
z up. A box rotation R = [ex ey ez]
has the diver's body axes as its columns, expressed in the sonar frame:
ex— torso forward (chest normal); extentlengthey— right shoulder to left shoulder; extentwidthez = ex × ey— feet to head; extentheight
So for a diver height is the head-to-toe long axis (median 1.74 m), whatever
direction the diver is facing. Rotations are exchanged as unit quaternions in Hamilton
(w, x, y, z) order; the scorer also accepts the continuous 6D representation or a
3×3 matrix. The dataset's FORMAT.md has the reference implementation.
Evaluation
Detections are matched to ground truth by exact volumetric IoU between arbitrarily oriented cuboids — not the rotated-BEV-times-height-overlap shortcut, which is invalid once a target tilts. From those matches the scorer reports:
- AP3D at IoU 0.30 / 0.35 / 0.40 / 0.50, all-point interpolated, and mAP3D = mean over {0.30, 0.35, 0.40}.
- APBEV@0.35 on the true bird's-eye footprint (a hexagon for a tilted diver). Diagnostic only.
- Over true positives at IoU ≥ 0.35: ATE (centre error, m), ASE
(1 − IoU after aligning centre and orientation), AOE (yaw only),
AOE3D (full-
SO(3)geodesic error) and tilt (angle between predicted and true body up-axes).
AOE3D is the metric that defines this challenge. It is minimised only over the
cuboid's own yaw symmetry {I, Rz(π)}; head-for-feet or front-for-back
is scored as the 180° error it is. A yaw-only detector can do well on APBEV and AOE and still
be badly wrong on AOE3D — in the organisers' ablation, AOE barely moves across detection heads
while AOE3D drops from 65–71° to 48° once the box is lifted to SO(3).
The leaderboard is ranked by the Sonar–Diver Detection Score, adapted from the nuScenes Detection Score:
SDS = ( 3 · mAP3D + (1 − min(1, ATE / 1 m)) + (1 − ASE) + (1 − min(1, AOE3D / 90°)) ) / 6
Half of the score is detection quality; the other half is split evenly across localisation, scale and
full-3D orientation. Ties are broken by AP3D@0.50, then AOE3D. The scorer is shipped inside the
dataset package as tools/diver3d_eval.py (NumPy + SciPy, CPU only) — the same file
the server runs — so a score on the validation split locally is the score you would get on the server.
Paper reference results
These figures are from the SonarVoxNet paper and have not yet been re-scored with the released challenge scorer. They are not verified challenge baselines or competing leaderboard entries. Updated reference results will follow after re-scoring.
| method | rotation | AP3D@0.35 | AP3D@0.50 | mAP3D | ATE (m) | ASE | AOE3D (°) | SDS |
|---|---|---|---|---|---|---|---|---|
| VoxelNet | yaw | 0.126 | 0.016 | – | – | – | – | 0.323 |
| PointPillars | yaw | 0.190 | 0.018 | – | – | – | – | 0.370 |
| SECOND | yaw | 0.315 | 0.016 | – | – | – | – | 0.436 |
| CenterPoint | yaw | 0.324 | 0.019 | – | – | – | – | 0.433 |
| VoxelNeXt | yaw | 0.374 | 0.019 | – | – | – | – | 0.463 |
| SonarVoxNet | SO(3) | 0.703 | 0.359 | 0.690 | 0.183 | 0.336 | 46.2 | 0.673 |
Two-seed means. The yaw-only detectors' AP3D@0.50 collapses because a yaw-only box cannot reach IoU 0.5 against a tilted diver; their geometry terms were not reported.
Submission format
One JSON file mapping each test sample_token ("<scene>/<frame>",
listed in annotations/test_tokens.txt) to its detections:
{
"results": {
"scene_0021/000000": [
{"center": [3.21, -0.44, 0.10],
"dimensions": [0.88, 0.86, 1.71],
"quaternion": [0.9848, 0.0, 0.1736, 0.0],
"score": 0.93}
],
"scene_0021/000001": []
}
}
scoreis required and must be in[0, 1]. AP depends on the ranking; a constant score will do badly.- At most 100 detections per frame are scored, kept by score. Do not pad.
- Frames you omit are scored as empty. Tokens not in the test split fail the submission.
- The same content as gzip-compressed JSON lines (
{"sample_token": ..., "boxes": [...]}per line) is also accepted and is recommended for large submissions; the format is detected from the content, so upload it with the.jsonextension. - Maximum upload 150 MB, measured both as uploaded and after decompression. A score-thresholded submission is typically under 20 MB.
- Run
tools/validate_submission.py --pred your.json --tokens annotations/test_tokens.txtbefore uploading. It catches coordinate-frame mistakes (e.g. a flipped axis puts every box outside the capture volume) in seconds.
Participate
- Download the dataset and read
FORMAT.md. - Train on
train; select checkpoints onvalwithtools/diver3d_eval.py --gt annotations/gt_val.json --pred your_val.json. - Run inference on every token in
annotations/test_tokens.txt. - Validate the file, then submit your predictions using a verified MaCVi account. Results appear on your dashboard after evaluation; queued submissions may take longer. Failed submissions count against the daily limit.
Rules
- Train on the train split only. Validation labels may be used for model selection and local scoring. Do not train on validation or test point clouds, including through self-supervised or test-time adaptation methods.
- External data, simulation and pre-training are allowed and must be disclosed at submission.
- Camera imagery from the recording rig is not released and may not be used.
- Organiser reference entries may appear on the leaderboard; they do not compete for a winning position.
Data provenance and citation
Diver3D is derived from uScenes, a multimodal RGB + 3D multibeam sonar dataset
from the same Blue Grotto recording campaign (110 scenes, 95,834 synchronised observations). The
Diver3D point clouds are uScenes sonar frames with full-SO(3) diver boxes added. The
dataset is released under CC BY-NC 4.0;
the tools are MIT. If you use Diver3D, please cite uScenes and the SonarVoxNet paper:
@inproceedings{dong2026uscenes,
title = {uScenes: A Multimodal {RGB} and {3D} Sonar Dataset for Underwater Robot Perception},
author = {Dong, Trung and Wu, Zhenqi and Penumarti, Aditya and Zhang, Zi-Hao
and Bartlett, Micaiah and Shin, Jane and Lin, Xiaomin},
booktitle = {OCEANS 2026 Monterey},
publisher = {IEEE/MTS},
year = {2026},
url = {https://github.com/era-research-lab/uScenes}
}
@article{park2026sonarvoxnet,
title = {SonarVoxNet: Diver Detection in {3D} Bounding Box using {3D} Sonar},
author = {Park, Eugene and Lee, Jiwon and Kan, Seyoung and Dong, Trung and Lin, Xiaomin and Shin, Jane},
journal = {arXiv preprint arXiv:2610.01644},
year = {2026},
url = {https://arxiv.org/abs/2610.01644}
}
The SonarVoxNet paper (arXiv:2610.01644) introduces the annotations, the detector and the evaluation protocol used here.
Updates and support
Release updates and questions: the MaCVi Discord community.
This challenge is hosted by the University of South Florida.














