5th Workshop on Maritime Computer Vision (MaCVi)

MaCVi @ WACV 2027

Challenges / Diver3D sonar detection

Diver3D: 3D Sonar Diver Detection

Detect divers as full 9-DoF oriented boxes from sparse 3D sonar point clouds.

An autonomous underwater vehicle assisting a human diver must track not just where the diver is, but how their body is oriented — and underwater, cameras fail first. Forward-looking sonar survives turbidity but projects the scene onto a 2D plane, discarding the elevation that orientation depends on. Recently commercialised 3D sonar keeps elevation, at the price of sparse, noisy, multipath-corrupted returns that dense-LiDAR detectors were never designed for. And unlike cars and pedestrians, a free-swimming diver pitches and rolls continuously: the yaw-only box used throughout automotive 3D detection cannot represent the pose at all.

Left to right: the camera view of a swimming diver; the same diver as 3D sonar returns; the annotated full-SO(3) box; and the best box a yaw-only detector could produce, which misses the tilted body.

The same diver, four ways. A yaw-only box with the correct centre and extents still cannot enclose a diver tilted 72° from vertical — that is the gap this challenge measures.

Diver3D v2 is available. 36 scenes, 37,010 sonar frames and 34,568 annotated divers, with labels for train and val and withheld test labels. Point clouds for all three splits are included, together with the official scorer. Online submissions are open. See the provisional challenge schedule.

Download Diver3D v2 (3.45 GB) SHA-256 checksum

Submit predictions Leaderboard

The link opens a shared folder containing diver3d_v2.zip and diver3d_v2.zip.sha256. Expected SHA-256 of the zip: 73bcee1251f5f12932859e4ef421e646c569124211628633989d4303aa5f89cb — verify with sha256sum -c diver3d_v2.zip.sha256.

Quick links: Task Dataset Box convention Evaluation Submission Ask for help

Task

Given one 3D sonar point cloud — an (N, 4) array of x, y, z, intensity in the sonar frame, with N between 428 and 11,965 — output every diver as a 9-DoF oriented box: a 3D centre, extents (length, width, height), a full rotation R ∈ SO(3), and a confidence score. There is one class (Diver); 27.6 % of frames contain no diver and are scored as real negatives. Frames are provided in acquisition order with stable instance ids on the training labels, so temporal methods are allowed, but the task is scored per frame.

Dataset

Diver3D was collected at Blue Grotto, a natural spring cave-diving site in Williston, Florida, across seven field sessions. Divers in a natural site exhibit the pitch and roll rarely seen in pool recordings: 34.8 % of annotated instances are near-upright, 30.7 % near-horizontal, and the rest in between. Sonar frames are annotated by hand with a full-SO(3) box enclosing the complete perceived body extent, using synchronised camera imagery as the annotation aid. The camera imagery is not part of the challenge release; this is a sonar-only task.

splitscenesframesannotated diverslabels
train2120,86119,482released
val78,0377,509released
test88,1127,577withheld
total3637,01034,568

The split is scene-wise, so temporally adjacent frames never straddle a split. Four divers appear in the data and every split contains at least three of them, so the same people are seen in train, val and test. That is a property of the dataset: results measure detection of diver bodies in this environment, not generalisation to unseen individuals, and should be cited as such. Point clouds span x ∈ [0.1, 16.0], y ∈ [−11.3, 11.3], z ∈ [−5.5, 5.5] m; the organisers' baseline crops to [0, 12] × [−5, 5] × [−2.5, 2.5] m, which keeps 99.9 % of annotated divers, but any crop is allowed and evaluation is over all annotated instances.

Box convention

This is where sonar 3D detection goes wrong most often, so it is stated precisely. The sonar frame is right-handed in metres: x forward along the boresight, y left, z up. A box rotation R = [ex ey ez] has the diver's body axes as its columns, expressed in the sonar frame:

So for a diver height is the head-to-toe long axis (median 1.74 m), whatever direction the diver is facing. Rotations are exchanged as unit quaternions in Hamilton (w, x, y, z) order; the scorer also accepts the continuous 6D representation or a 3×3 matrix. The dataset's FORMAT.md has the reference implementation.

Evaluation

Detections are matched to ground truth by exact volumetric IoU between arbitrarily oriented cuboids — not the rotated-BEV-times-height-overlap shortcut, which is invalid once a target tilts. From those matches the scorer reports:

AOE3D is the metric that defines this challenge. It is minimised only over the cuboid's own yaw symmetry {I, Rz(π)}; head-for-feet or front-for-back is scored as the 180° error it is. A yaw-only detector can do well on APBEV and AOE and still be badly wrong on AOE3D — in the organisers' ablation, AOE barely moves across detection heads while AOE3D drops from 65–71° to 48° once the box is lifted to SO(3).

The leaderboard is ranked by the Sonar–Diver Detection Score, adapted from the nuScenes Detection Score:

SDS = ( 3 · mAP3D + (1 − min(1, ATE / 1 m)) + (1 − ASE) + (1 − min(1, AOE3D / 90°)) ) / 6

Half of the score is detection quality; the other half is split evenly across localisation, scale and full-3D orientation. Ties are broken by AP3D@0.50, then AOE3D. The scorer is shipped inside the dataset package as tools/diver3d_eval.py (NumPy + SciPy, CPU only) — the same file the server runs — so a score on the validation split locally is the score you would get on the server.

Paper reference results

These figures are from the SonarVoxNet paper and have not yet been re-scored with the released challenge scorer. They are not verified challenge baselines or competing leaderboard entries. Updated reference results will follow after re-scoring.

methodrotationAP3D@0.35AP3D@0.50mAP3DATE (m)ASEAOE3D (°)SDS
VoxelNetyaw0.1260.016––––0.323
PointPillarsyaw0.1900.018––––0.370
SECONDyaw0.3150.016––––0.436
CenterPointyaw0.3240.019––––0.433
VoxelNeXtyaw0.3740.019––––0.463
SonarVoxNetSO(3)0.7030.3590.6900.1830.33646.20.673

Two-seed means. The yaw-only detectors' AP3D@0.50 collapses because a yaw-only box cannot reach IoU 0.5 against a tilted diver; their geometry terms were not reported.

Submission format

One JSON file mapping each test sample_token ("<scene>/<frame>", listed in annotations/test_tokens.txt) to its detections:

{
  "results": {
    "scene_0021/000000": [
      {"center": [3.21, -0.44, 0.10],
       "dimensions": [0.88, 0.86, 1.71],
       "quaternion": [0.9848, 0.0, 0.1736, 0.0],
       "score": 0.93}
    ],
    "scene_0021/000001": []
  }
}

Participate

  1. Download the dataset and read FORMAT.md.
  2. Train on train; select checkpoints on val with tools/diver3d_eval.py --gt annotations/gt_val.json --pred your_val.json.
  3. Run inference on every token in annotations/test_tokens.txt.
  4. Validate the file, then submit your predictions using a verified MaCVi account. Results appear on your dashboard after evaluation; queued submissions may take longer. Failed submissions count against the daily limit.

Rules

Data provenance and citation

Diver3D is derived from uScenes, a multimodal RGB + 3D multibeam sonar dataset from the same Blue Grotto recording campaign (110 scenes, 95,834 synchronised observations). The Diver3D point clouds are uScenes sonar frames with full-SO(3) diver boxes added. The dataset is released under CC BY-NC 4.0; the tools are MIT. If you use Diver3D, please cite uScenes and the SonarVoxNet paper:

@inproceedings{dong2026uscenes,
  title     = {uScenes: A Multimodal {RGB} and {3D} Sonar Dataset for Underwater Robot Perception},
  author    = {Dong, Trung and Wu, Zhenqi and Penumarti, Aditya and Zhang, Zi-Hao
               and Bartlett, Micaiah and Shin, Jane and Lin, Xiaomin},
  booktitle = {OCEANS 2026 Monterey},
  publisher = {IEEE/MTS},
  year      = {2026},
  url       = {https://github.com/era-research-lab/uScenes}
}
@article{park2026sonarvoxnet,
  title   = {SonarVoxNet: Diver Detection in {3D} Bounding Box using {3D} Sonar},
  author  = {Park, Eugene and Lee, Jiwon and Kan, Seyoung and Dong, Trung and Lin, Xiaomin and Shin, Jane},
  journal = {arXiv preprint arXiv:2610.01644},
  year    = {2026},
  url     = {https://arxiv.org/abs/2610.01644}
}

The SonarVoxNet paper (arXiv:2610.01644) introduces the annotations, the detector and the evaluation protocol used here.

Updates and support

Release updates and questions: the MaCVi Discord community.


This challenge is hosted by the University of South Florida.