Expand same-scale benchmark comparisons

#2
by leyili6666 - opened
Files changed (1) hide show
  1. README.md +3 -0
README.md CHANGED
@@ -72,6 +72,9 @@ Scores (%); evaluation settings vary by source.
72
  | GSM8K | Qwen2-72B-Instruct | 72B | 93.2 | **93.75** | [Qwen2.5 report, Table 6](https://arxiv.org/html/2412.15115v2#S5.SS2.SSS1) |
73
  | GSM8K | SciTulu-70B | 70B | 67.5 | **93.75** | [SciRIFF report, Table 7](https://arxiv.org/html/2406.07835v2#A3) |
74
  | GSM8K | WizardMath-Llama-RL (Llama 2) | 70B | 92.8 | **93.75** | [WizardMath report, Tables 1 & 15](https://arxiv.org/html/2308.09583v2) |
 
 
 
75
  | ARC-Easy | DeepSeek-LLM-67B-Chat | 67B | 81.6 | **84.64** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
76
  | ARC-Challenge | DeepSeek-LLM-67B-Chat | 67B | 64.1 | **64.42** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
77
 
 
72
  | GSM8K | Qwen2-72B-Instruct | 72B | 93.2 | **93.75** | [Qwen2.5 report, Table 6](https://arxiv.org/html/2412.15115v2#S5.SS2.SSS1) |
73
  | GSM8K | SciTulu-70B | 70B | 67.5 | **93.75** | [SciRIFF report, Table 7](https://arxiv.org/html/2406.07835v2#A3) |
74
  | GSM8K | WizardMath-Llama-RL (Llama 2) | 70B | 92.8 | **93.75** | [WizardMath report, Tables 1 & 15](https://arxiv.org/html/2308.09583v2) |
75
+ | AQuA-RAT | Llama-2-70B-Chat | 70B | 31.32 | **77.56** | [Diversity of Thought paper](https://openreview.net/pdf?id=FvfhHucpLd) |
76
+ | ARC-Easy | Llama-2-70B | 70B | 76.5 | **84.64** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
77
+ | ARC-Challenge | Llama-2-70B | 70B | 59.5 | **64.42** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
78
  | ARC-Easy | DeepSeek-LLM-67B-Chat | 67B | 81.6 | **84.64** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
79
  | ARC-Challenge | DeepSeek-LLM-67B-Chat | 67B | 64.1 | **64.42** | [DeepSeek official results](https://github.com/deepseek-ai/DeepSeek-LLM/blob/main/evaluation/more_results.md) |
80