Summary of the bundled anonymized manuscript; findings refer to that study’s own comparisons.
Question
How useful is cortical geometry without fMRI-specific encoder training?
Approach
Render surface activity as flatmap sequences and apply a frozen SigLIP2 encoder.
Evidence
Resting-state tasks including HCP, ADNI, and PPMI; visual decoding with NSD.
Finding
The bundled draft reports gains over ROI baselines on HCP and ADNI, weaker results on PPMI, and performance below the strongest voxel models overall.
Source and scope
This fMRI Atlas page is a curated research summary, not the original study. The question, method, evidence and finding above summarize the authors’ report; results have not been independently reproduced here. See the bundled anonymized manuscript for methods, authorship and full results.
Evidence boundary: The bundled anonymized manuscript reports mixed results, including weaker PPMI performance and results below stronger voxel baselines; its public bibliographic record should be checked before reuse.
Page reviewed: .
Abstract
Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized 3D/4D architectures and costly fMRI-specific pretraining. We ask whether part of this performance gap reflects the importance of preserving cortical geometry. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a frozen-encoder surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Across resting-state benchmarks, FlatClip serves as a competitive middle-ground representation, outperforming ROI-level baselines on HCP and ADNI tasks while remaining weaker on PPMI and below the strongest voxel-level models overall. On visual-fMRI decoding, performance improves when the input is restricted from whole cortex to visual or NSD-provided task-active cortex, suggesting that geometry is most useful when it is aligned with the prediction target. We further probe this interpretation with geometry controls: disrupting flatmap spatial organization reduces performance, mapping ROI-level signals back into atlas-defined 4D volumes partially recovers performance, and adaptive patching in highly activated regions improves several voxel-based prediction metrics. Together, these results position surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models, and suggest that geometry-preserving spatial organization can be useful when designing fMRI models. Code is available.
Figure 1 · Spatial granularity, from ROI to surface and voxel representations.
Method
FlatClip renders cortical fMRI activity as geometry-aware flatmaps and extracts image features using a frozen SigLIP2 encoder. Only a lightweight downstream probe is trained.
Resting-state fMRI
Each subject is represented by 40 flatmap frames. Frame-level features are averaged over time to produce a single 768-dimensional subject representation for downstream prediction.
Visual-task fMRI
Each stimulus-level GLM response map is rendered as a flatmap and encoded into one feature vector. Experiments compare whole-cortex, visual-cortex, and NSDgeneral task-active regions.
Source status
FlatClip is accepted to NeurIPS 2026. The bundled PDF remains anonymized, and a public author list, code URL, and official proceedings citation are not supplied in this package; add the final BibTeX once those details are available.