FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models
Abstract
Remote sensing foundation models (RSFMs) are commonly evaluated using aggregate metrics, which can hide systematic performance disparities across ecological regions. We introduce FairRSFM, a biome-aware benchmark for evaluating ecological group robustness in RSFMs. FairRSFM maps georeferenced samples from 14 terrestrial biome classes into six ecologically meaningful macro-groups and evaluates models under a unified frozen-backbone evaluation protocol. The benchmark covers four downstream datasets: m-EuroSAT, m-BigEarthNet, m-SA-Crop-Type, and MMEarth20K with Dynamic World label maps. Using Prithvi-EO-2.0, SatMAE, and DOFA across three random seeds, we show that aggregate performance consistently masks biome-dependent disparities across architectures and tasks. For example, Prithvi-EO-2.0 reaches 90.98% overall macro-F1 on m-EuroSAT but a mean worst-group score of only 83.72%, while m-SA-Crop-Type drops from 27.30% overall mIoU to 18.47% in the Xeric and Mineralogical group. We further evaluate Biome-Orthogonal Linear Probing (BOLP), Dynamic Biome Reweighting (DBR), and GroupDRO as complementary mitigation baselines. Their effectiveness is model- and task-dependent; for example, BOLP improves Prithvi-EO-2.0 worst-group F1@opt on m-BigEarthNet from 46.12% to 50.27% without updating the RSFM backbone. FairRSFM provides a reusable protocol for diagnosing and mitigating ecological robustness gaps in remote sensing foundation models. Code and datasets are available at: https://github.com/aminurhossain/FairRSFM.
Community
We introduce FairRSFM, a biome-aware benchmark and debiasing framework for evaluating ecological robustness in Remote Sensing Foundation Models (RSFMs).
Our results show that strong aggregate performance can mask substantial performance disparities across ecological regions. FairRSFM provides a unified framework for biome-aware evaluation and investigates BOLP, Dynamic Biome Reweighting (DBR), and GroupDRO as mitigation strategies.
Code and datasets are available with the paper. We welcome feedback, discussions, and collaborations!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RSPDBench: Benchmarking Vision Foundation Models on Earth Observation Tasks Under Physically Grounded Remote-Sensing Product Degradations (2026)
- Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling (2026)
- Cryo-Bench: Benchmarking Foundation Models for Cryosphere Mapping (2026)
- Foundation Models Meet Agriculture: Challenges Beyond Pretraining (2026)
- Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts (2026)
- SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping (2026)
- Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.05790 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper