Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
AI & ML interests
Factuality, reasoning, alignment, LLM applications
Recent Activity
View all activity
Papers
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
spaces 7
Running
LudoBench
🎲
Multimodal Game Reasoning Benchmark [ICLR 2026]
Sleeping
Agents
Answer Convergence Early Stopping
🛑
Demo for EMNLP Paper "Answer Convergence as a Signal..."
Running
FactRBench
🏆
View and analyze long-form factuality leaderboard
Sleeping
3
ExpertLongBench
🚀
Leaderboard for ExpertLongBench
Sleeping
1
ManyICLBench
🚀
Leaderboard for ManyICLBench
Running
MLRC-BENCH
📊
Display model performance rankings
models 15
launch/MET-D-Gemma3-4B-en-only
Text Generation • 4B • Updated • 130
launch/MET-D-Gemma3-4B
Text Generation • 4B • Updated • 125
launch/MET-D-Qwen3-8B-en-only
Text Generation • 8B • Updated • 118
launch/MET-D-Qwen3-8B
Text Generation • 8B • Updated • 143
launch/MET-D-Qwen3-4B-zh-only
Text Generation • 4B • Updated • 140
launch/MET-D-Qwen3-4B-ms-only
Text Generation • 4B • Updated • 141
launch/MET-D-Qwen3-4B-ko-only
Text Generation • 4B • Updated • 133
launch/MET-D-Qwen3-4B-hi-only
Text Generation • 4B • Updated • 108
launch/MET-D-Qwen3-4B-es-only
Text Generation • 4B • Updated • 136
launch/MET-D-Qwen3-4B-en-only
Text Generation • 4B • Updated • 125
datasets 14
launch/MCLASH
Viewer • Updated • 2.61k • 176
launch/CLASH
Viewer • Updated • 345 • 245 • 3
launch/thinkprm-1K-verification-cots
Viewer • Updated • 1k • 106 • 8
launch/LudoBench
Viewer • Updated • 638 • 71
launch/ExpertLongBench
Preview • Updated • 342 • 10
launch/ManyICLBench
Viewer • Updated • 66 • 264 • 1
launch/CMV
Viewer • Updated • 133 • 15
launch/FactRBench
Viewer • Updated • 1.06k • 35 • 2
launch/FactBench
Viewer • Updated • 1k • 107 • 3
launch/gov_report
Viewer • Updated • 58.4k • 885 • 14