AI & ML interests

Frontier alignment research to ensure the safe development and deployment of advanced AI systems.

Recent Activity

skar0  updated a dataset 9 days ago
AlignmentResearch/impossible-swegym
skar0  published a dataset 14 days ago
AlignmentResearch/impossible-swegym
sam-far  published a model about 1 month ago
AlignmentResearch/sam266-gemma-3-27b-merged
View all activity

AlignmentResearch 's collections 4

The Obfuscation Atlas
Obfuscated Policy, Obfuscated Activations, Blatant Deception, and Honest models trained in the Obfuscation Atlas paper.
The Obfuscation Altas
Obfuscated Policy, Obfuscated Activations, Blatant Deception, and Honest models trained in the Obfuscation Atlas paper