SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
Paper • 2602.14687 • Published
How to use decoderesearch/synth-sae-bench-16k-v1-saes with SAELens:
# pip install sae-lens
from sae_lens import SAE
sae, cfg_dict, sparsity = SAE.from_pretrained(
release = "RELEASE_ID", # e.g., "gpt2-small-res-jb". See other options in https://github.com/jbloomAus/SAELens/blob/main/sae_lens/pretrained_saes.yaml
sae_id = "SAE_ID", # e.g., "blocks.8.hook_resid_pre". Won't always be a hook point
)Sample Sparse Autoencoders (SAEs) trained on the SynthSAEBench-16k-v1 model. Training code is at https://github.com/decoderesearch/synth-sae-bench-experiments.
We train 5 different SAE types, each with 5 seeds and L0 from 15-45. Each SAE has a stats.json file containing eval stats.
Check out SAELens to train your own SAEs on SynthSAEBench and make your own custom synthetic data models. Also see the SynthSAEBench paper for more details.