Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Mechanist Interpretability for Alignment Algorithms
community
Activity Feed
Follow
5
AI & ML interests
AI Safety, Mechanist Interpretability
Team members
5
MInAlA
's models
18
Sort: Recently updated
MInAlA/Llama-3.2-3B-Instruct-PPO-merged
Text Generation
•
3B
•
Updated
Apr 23
•
1.25k
MInAlA/SmolLM3-3B-PPO-merged
3B
•
Updated
Apr 22
•
467
MInAlA/Qwen3-4B-Instruct-2507-PPO-merged
Text Generation
•
4B
•
Updated
Apr 20
•
832
MInAlA/Llama-3.2-3B-SimPO-merged
Text Generation
•
3B
•
Updated
Apr 18
•
1.28k
MInAlA/Qwen3-4B-Instruct-2507-SimPO-merged
Text Generation
•
4B
•
Updated
Apr 18
•
865
MInAlA/SmolLM3-3B-SimPO-merged
Text Generation
•
3B
•
Updated
Apr 17
•
544
MInAlA/Llama-3.2-3B-Instruct-GRPO-merged
Text Generation
•
3B
•
Updated
Apr 16
•
1.2k
MInAlA/Qwen3-4B-Instruct-2507-GRPO-merged
Text Generation
•
4B
•
Updated
Apr 14
•
863
MInAlA/SmolLM3-3B-GRPO-merged
Text Generation
•
3B
•
Updated
Apr 12
•
551
MInAlA/Llama-3.2-3B-Instruct-KTO-merged
Text Generation
•
3B
•
Updated
Apr 11
•
1.29k
MInAlA/Qwen3-4B-Instruct-2507-KTO-merged
Text Generation
•
4B
•
Updated
Apr 11
•
862
MInAlA/Qwen3-4B-ORPO-merged
4B
•
Updated
Apr 10
•
805
MInAlA/Llama-3.2-3B-ORPO-merged
Text Generation
•
3B
•
Updated
Apr 10
•
1.32k
MInAlA/SmolLM3-3B-KTO-merged
Text Generation
•
3B
•
Updated
Apr 10
•
556
MInAlA/SmolLM3-3B-ORPO-merged
Text Generation
•
3B
•
Updated
Apr 6
•
559
MInAlA/Llama-3.2-3B-DPO-merged
Text Generation
•
3B
•
Updated
Apr 5
•
1.36k
MInAlA/Qwen3-4B-Instruct-2507-DPO-merged
Text Generation
•
4B
•
Updated
Apr 5
•
899
MInAlA/SmolLM3-3B-DPO-merged
Text Generation
•
3B
•
Updated
Apr 5
•
569