WW2Bench Leaderboard
WW2Bench scores for small language models, v0.1.0.
I see you mentioned in the README using gpt and claude. Have you determined this is ok under anthropic's terms? I suppose it is because Anthropic offers no model training framework and therefore it's not competing ?? Not sure on this.
I read the chat - this is really entertaining. I need to check this out.