Instructions to use phanerozoic/argus-3d with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- EUPE
How to use phanerozoic/argus-3d with EUPE:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
SAM3 negative + ensemble negative + mAP vs Boxer 41.2 / CuTR 25.0 logged
Browse files
TODO.md
CHANGED
|
@@ -34,6 +34,8 @@
|
|
| 34 |
|
| 35 |
17. Iterative box refinement (full). After initial fusion, project the world-frame fused boxes back to camera frame as priors. Re-fit per-frame OBBs constrained to the inlier patches inside an enlarged version of the projected prior; re-fuse. This requires keeping the per-frame point clouds across passes (a real refactor); the current iterative refinement only does observation-level trimming, not point-level refit.
|
| 36 |
|
| 37 |
-
18. SAM
|
|
|
|
|
|
|
| 38 |
|
| 39 |
19. Multi-resolution head retrain. The 1024-input experiment regressed because the MLP head was trained at 768. Re-extracting features at 1024 from 20 scenes and retraining the MLP on 1024 features (4096 patches/frame instead of 2304) would let small objects (5 cm photo frames at 2 m depth, currently 1-2 patches) become 2-3 patches and potentially detectable.
|
|
|
|
| 34 |
|
| 35 |
17. Iterative box refinement (full). After initial fusion, project the world-frame fused boxes back to camera frame as priors. Re-fit per-frame OBBs constrained to the inlier patches inside an enlarged version of the projected prior; re-fuse. This requires keeping the per-frame point clouds across passes (a real refactor); the current iterative refinement only does observation-level trimming, not point-level refit.
|
| 36 |
|
| 37 |
+
18. SAM 3 prompt-based instance discovery. Tested across two iterations. First pass: SAM 3 (`facebook/sam3`, 840 M params) prompted with 22 indoor-class concepts per frame, masks unprojected to 3D via GT depth, OBB fit on the unprojected points. Result: mean IoU 0.028, recall 30 %. Second pass, ensemble (CC observations + SAM 3 observations into one IoU-fusion pool): mean IoU 0.157 with SAM 3 weight 0.1 / 0.147 with weight 0.5, recall 27-29 %. Both ensembles buy ~10 percentage points of recall but cost ~6 mean-IoU points vs CC-only (0.216 / 17.5 %). The IoU regression is structural: SAM 3 masks tightly cover the visible surface only (legs of a chair excluded if hidden), so the per-frame OBB depth dimension is underestimated, and even at low weight in the multi-view fusion the SAM 3 observations join CC clusters and pull the weighted-median fuse off-axis. SAM 3 is a useful recall-boost component but not a drop-in replacement for the CC-based instance discovery. The credible next iteration is to use SAM 3 only as an instance separator: keep its 2D mask, but fit the OBB on the CC pipeline's full per-component point cloud restricted to that mask's 2D footprint, never on the SAM 3 mask pixels themselves. That avoids partial-view OBB collapse while keeping SAM 3's diversity.
|
| 38 |
+
|
| 39 |
+
19. mAP scoring across IoU thresholds [0.05, 0.10, 0.15, 0.25, 0.50] (TODO 10). Done. Strict default (44 fused boxes): AP @ 0.05 = 14.4, AP @ 0.50 = 0.4, mAP = **7.5**. Loose filter (246 fused boxes): AP @ 0.05 = 24.6, AP @ 0.50 = 0.1, mAP = **11.9**. Boxer reports **41.2 mAP** on CA-1M with dense depth + GT 2D bboxes + a learned 2D→3D lifter (vs CuTR's 25.0 mAP on the same setting). Our 11.9 is at ~29 % of Boxer's mAP and ~48 % of CuTR's, while running class-agnostic without GT 2D bboxes and with a closed-form OBB fitter. The ~4× gap at AP @ 0.5 IoU is dominated by the percentile-OBB fit on point clouds; the field's range from "almost nothing works" to "with every aid it works" is roughly 1 mAP (CuTR egocentric without dense depth) to 53 mAP (Boxer egocentric without dense depth), so 12 mAP class-agnostic-no-GT-2D is a reasonable mid-range result for the configuration.
|
| 40 |
|
| 41 |
19. Multi-resolution head retrain. The 1024-input experiment regressed because the MLP head was trained at 768. Re-extracting features at 1024 from 20 scenes and retraining the MLP on 1024 features (4096 patches/frame instead of 2304) would let small objects (5 cm photo frames at 2 m depth, currently 1-2 patches) become 2-3 patches and potentially detectable.
|