phanerozoic commited on
Commit
de2ae5b
·
verified ·
1 Parent(s): 6f137fe

SAM3 negative + ensemble negative + mAP vs Boxer 41.2 / CuTR 25.0 logged

Browse files
Files changed (1) hide show
  1. TODO.md +3 -1
TODO.md CHANGED
@@ -34,6 +34,8 @@
34
 
35
  17. Iterative box refinement (full). After initial fusion, project the world-frame fused boxes back to camera frame as priors. Re-fit per-frame OBBs constrained to the inlier patches inside an enlarged version of the projected prior; re-fuse. This requires keeping the per-frame point clouds across passes (a real refactor); the current iterative refinement only does observation-level trimming, not point-level refit.
36
 
37
- 18. SAM-style prompt-based instance discovery. The per-frame instance separation is currently CC-on-fg-mask + DBSCAN. SAM-2 or DINOv3 mask-from-point prompting could replace it: prompt with the highest-fg patch, get a clean instance mask, fit OBB on its 3D unprojected points. Not zero-parameter anymore, but might break the 0.177 mean IoU ceiling.
 
 
38
 
39
  19. Multi-resolution head retrain. The 1024-input experiment regressed because the MLP head was trained at 768. Re-extracting features at 1024 from 20 scenes and retraining the MLP on 1024 features (4096 patches/frame instead of 2304) would let small objects (5 cm photo frames at 2 m depth, currently 1-2 patches) become 2-3 patches and potentially detectable.
 
34
 
35
  17. Iterative box refinement (full). After initial fusion, project the world-frame fused boxes back to camera frame as priors. Re-fit per-frame OBBs constrained to the inlier patches inside an enlarged version of the projected prior; re-fuse. This requires keeping the per-frame point clouds across passes (a real refactor); the current iterative refinement only does observation-level trimming, not point-level refit.
36
 
37
+ 18. SAM 3 prompt-based instance discovery. Tested across two iterations. First pass: SAM 3 (`facebook/sam3`, 840 M params) prompted with 22 indoor-class concepts per frame, masks unprojected to 3D via GT depth, OBB fit on the unprojected points. Result: mean IoU 0.028, recall 30 %. Second pass, ensemble (CC observations + SAM 3 observations into one IoU-fusion pool): mean IoU 0.157 with SAM 3 weight 0.1 / 0.147 with weight 0.5, recall 27-29 %. Both ensembles buy ~10 percentage points of recall but cost ~6 mean-IoU points vs CC-only (0.216 / 17.5 %). The IoU regression is structural: SAM 3 masks tightly cover the visible surface only (legs of a chair excluded if hidden), so the per-frame OBB depth dimension is underestimated, and even at low weight in the multi-view fusion the SAM 3 observations join CC clusters and pull the weighted-median fuse off-axis. SAM 3 is a useful recall-boost component but not a drop-in replacement for the CC-based instance discovery. The credible next iteration is to use SAM 3 only as an instance separator: keep its 2D mask, but fit the OBB on the CC pipeline's full per-component point cloud restricted to that mask's 2D footprint, never on the SAM 3 mask pixels themselves. That avoids partial-view OBB collapse while keeping SAM 3's diversity.
38
+
39
+ 19. mAP scoring across IoU thresholds [0.05, 0.10, 0.15, 0.25, 0.50] (TODO 10). Done. Strict default (44 fused boxes): AP @ 0.05 = 14.4, AP @ 0.50 = 0.4, mAP = **7.5**. Loose filter (246 fused boxes): AP @ 0.05 = 24.6, AP @ 0.50 = 0.1, mAP = **11.9**. Boxer reports **41.2 mAP** on CA-1M with dense depth + GT 2D bboxes + a learned 2D→3D lifter (vs CuTR's 25.0 mAP on the same setting). Our 11.9 is at ~29 % of Boxer's mAP and ~48 % of CuTR's, while running class-agnostic without GT 2D bboxes and with a closed-form OBB fitter. The ~4× gap at AP @ 0.5 IoU is dominated by the percentile-OBB fit on point clouds; the field's range from "almost nothing works" to "with every aid it works" is roughly 1 mAP (CuTR egocentric without dense depth) to 53 mAP (Boxer egocentric without dense depth), so 12 mAP class-agnostic-no-GT-2D is a reasonable mid-range result for the configuration.
40
 
41
  19. Multi-resolution head retrain. The 1024-input experiment regressed because the MLP head was trained at 768. Re-extracting features at 1024 from 20 scenes and retraining the MLP on 1024 features (4096 patches/frame instead of 2304) would let small objects (5 cm photo frames at 2 m depth, currently 1-2 patches) become 2-3 patches and potentially detectable.