Skip to content
jndengPublic

About

[ECCV 2026] Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

Resources

Stars

32 stars

Watchers

1 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

Repository files navigation

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

Paper

SAF3R is a training-free dynamic sparse attention framework that accelerates existing feed-forward 3D reconstruction models, such as VGGT, by reducing the computational cost of global attention. It exploits head-wise sparsity heterogeneity to replace each full global attention head with the sparse attention kernel that best matches its attention pattern.

Teaser

What's in this repo

  • Visualization tools for analyzing head-wise global attention patterns in feed-forward 3D reconstruction models. Currently supported models include three offline models (VGGT, Pi3, and DA3) and one streaming model (StreamVGGT).
  • SAF3R offline profiling code that automatically assigns each global attention head to one of four predefined sparse attention patterns.
  • SAF3R inference patches that enable dynamic sparse attention inference for VGGT, Pi3, and DA3.
  • Benchmark and evaluation tools for evaluating model performance and efficiency. Currently supported benchmarks include DA3-Bench, Co3D-v2, RealEstate10K, and ScanNet.

Table of Contents

Installation

conda create -n saf3r python=3.10 -y
conda activate saf3r
git clone https://github.com/jndeng/SAF3R
cd SAF3R
pip install -e .

Analysis Tools

We provide visualization and analysis tools for frame/global attention patterns for each supported model under tools/. For each model, the corresponding tool is implemented as a standalone Jupyter notebook and can be run independently. The two main use cases are shown below.

Attention Distribution Overlay    Attention Maps of Layer Heads
Left: Semi-interactive visualization of the attention distribution for each selected query token (blue box), overlaid on the image. Right: Attention maps of all heads in a specific layer.

Model Inference

We provide inference example code for using SAF3R on different 3R models.

Note

Checkpoints will be automatically downloaded to the local cache directory checkpoints/ during the first run. They can also be manually downloaded from VGGT, Pi3, DA3-GIANT, and StreamVGGT.

Running Inference Demo

To run a demo inference script using a specific model on your image directories:

python scripts/inference_demo.py --model vggt --data_dir data/courthouse

Supported options for --model are vggt, pi3, and da3. By default, the predicted point clouds will be exported under tmp_plots/ as .ply files and browser-viewable interactive .html visualizations.

Note

The first run may take longer due to JIT compilation of the custom Triton kernels. The compiled kernels are cached and reused in subsequent runs.

Vis
Example interactive HTML visualization results.

Minimal Code Snippet (for SAF3R-VGGT)

Show code
import torch
from addict import Dict
from saf3r.utils.model_utils import build_model, infer_model

# 1. Configure the model and SAF3R patch settings
model_cfg = Dict(
    name="VGGT", 
    dpt_only=True, 
    ckpt_path="checkpoints/vggt/model_tracker_fixed_e20.pt",
    sparse_config_path="configs/sparse_attn/vggt/eth3d-train-fltr-calib_cmpmse.json",
    patch_module=Dict(
        type="headsparse",
        topk_mode="token",
        lazy_dino_topk=True,
        topk=4
    )
)

# 2. Build and automatically patch the model with SAF3R kernels
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = build_model(model_cfg, device)

# 3. Setup input data (image file paths)
scene_data = Dict(
    image_files=[
        "data/courthouse/000000.png",
        "data/courthouse/000001.png",
        ...
    ]
)

# 4. Perform dynamic sparse attention inference
pred_data, stats = infer_model(model, model_cfg, scene_data)

# Extract unified predictions
extrinsics = pred_data.extrinsics     # [N, 3, 4]
intrinsics = pred_data.intrinsics     # [N, 3, 3]
depth      = pred_data.depth          # [N, H, W]
depth_conf = pred_data.conf           # [N, H, W]

Evaluation Benchmarks

We provide evaluation code for multiple tasks and benchmarks.

Supported Benchmarks & Datasets

  • DA3-Bench Datasets (7Scenes, ETH3D, ScanNet++, HiRoom, DTU64, DTU)
    • Camera pose estimation
    • Video depth estimation
    • 3D point-cloud reconstruction
  • Co3D-v2
    • Camera pose estimation
  • RealEstate10K
    • Camera pose estimation
  • ScanNet (v2)
    • Camera pose estimation
    • 3D point-cloud reconstruction

Datasets Preparation

Please follow the corresponding instructions to prepare each dataset.

  • DA3-Bench (7Scenes, ETH3D, ScanNet++, HiRoom, DTU64, DTU)
  • Co3D (v2)
  • RealEstate10K
  • ScanNet (v2)
    • Follow the ScanNet instructions to download the dataset and place it under workspace/benchmark_dataset/. The list of the 50 evaluation scenes can be found here.

The downloaded datasets should be organized under workspace/benchmark_dataset/ as follows:

Show dataset structure
workspace/benchmark_dataset/
├── 7scenes/
│   └── 7Scenes/
│       ├── chess/
│       └── ...
├── eth3d/
│   ├── courtyard/
│   ├── electro/
│   └── ...
├── scannetpp/
│   ├── 09c1414f1b/
│   └── ...
├── hiroom/
│   ├── data/
│   ├── fused_pcd/
│   └── selected_scene_list_val.txt
├── dtu/
│   ├── Rectified/
│   ├── Cameras/
│   ├── Points/
│   ├── SampleSet/
│   └── depth_raw/
├── dtu64/
│   ├── Cameras/
│   ├── scan105/
│   └── ...
├── co3dv2/
│   ├── vggt_testset
│   │   ├── apple/
│   │   └── ...
│   └── vggt_anno
│       ├── apple_test.jgz
│       └── ...
├── realestate10k/
│   ├── data/re10k
│   │   ├── 005dd9a58df1ba3c/
│   │   └── ...
│   └── datasets/seq-id-maps
│       └── Re10K_relpose_seq-id-map_seed42.json
└── scannetv2/
    ├── scans/
    │   ├── scene0000_00/
    │   ├── scene0013_02/
    │   └── ...
    └── scannet_50.yaml

Running Evaluation

Run the unified launcher with the desired configuration file:

bash scripts/evaluate.sh eval_saf3r_vggt

Benchmarking Efficiency

To measure inference latency and peak memory usage across different sequences:

bash scripts/benchmark_efficiency.sh eval_saf3r_vggt 300

This runs the efficiency benchmark using the specified configuration file and sequence length.

Profiling Attention Heads

We provide profiled global-attention head configurations under configs/sparse_attn/.

To generate these profiling results from scratch:

  1. Download the calibration dataset (e.g., ETH3D) following the instructions in DA3-bench.
  2. Run the profiling script using the desired configuration file under configs/profile/:
    bash scripts/profile.sh profile_saf3r_vggt

Acknowledgements

This repository builds upon several excellent open-source projects, including VGGT, Pi3, DA3, StreamVGGT, LingBot-Map, FastVGGT, SparseVGGT, and Speed3R. We sincerely thank the authors and contributors for making their code publicly available.

Citation

If you find SAF3R useful for your research or project, please consider citing:

@article{deng2026saf3r,
  title={SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers},
  author={Deng, Jianing and Li, Yuanzhe and Wang, Jialu and Wang, Song and Chen, Tianlong and Yang, Huanrui and Hu, Jingtong},
  journal={arXiv preprint arXiv:2607.03612},
  year={2026}
}

About

[ECCV 2026] Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

Resources

Stars

32 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages