dongbobo commited on
Commit
b0ba55f
·
verified ·
1 Parent(s): f9e1f0b

Upload MyAwesomeModel (checkpoint step_1000): weights, config, figures and evaluation assets

Browse files
.gitattributes CHANGED
@@ -37,3 +37,4 @@ evaluation/build/lib.linux-x86_64-cpython-313/utils/benchmark_utils.cpython-313-
37
  evaluation/build/temp.linux-x86_64-cpython-313/utils/benchmark_utils.o filter=lfs diff=lfs merge=lfs -text
38
  evaluation/utils/benchmark_utils.cpython-313-x86_64-linux-gnu.so filter=lfs diff=lfs merge=lfs -text
39
  evaluation/utils/benchmark_utils.cpython-310-x86_64-linux-gnu.so filter=lfs diff=lfs merge=lfs -text
 
 
37
  evaluation/build/temp.linux-x86_64-cpython-313/utils/benchmark_utils.o filter=lfs diff=lfs merge=lfs -text
38
  evaluation/utils/benchmark_utils.cpython-313-x86_64-linux-gnu.so filter=lfs diff=lfs merge=lfs -text
39
  evaluation/utils/benchmark_utils.cpython-310-x86_64-linux-gnu.so filter=lfs diff=lfs merge=lfs -text
40
+ evaluation/utils/benchmark_utils.cpython-312-x86_64-linux-gnu.so filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -38,21 +38,21 @@ Beyond its improved reasoning capabilities, this version also offers a reduced h
38
 
39
  | | Benchmark | Model1 | Model2 | Model1-v2 | MyAwesomeModel |
40
  |---|---|---|---|---|---|
41
- | **Core Reasoning Tasks** | Math Reasoning | 0.510 | 0.535 | 0.521 | {RESULT} |
42
- | | Logical Reasoning | 0.789 | 0.801 | 0.810 | {RESULT} |
43
- | | Common Sense | 0.716 | 0.702 | 0.725 | {RESULT} |
44
- | **Language Understanding** | Reading Comprehension | 0.671 | 0.685 | 0.690 | {RESULT} |
45
- | | Question Answering | 0.582 | 0.599 | 0.601 | {RESULT} |
46
- | | Text Classification | 0.803 | 0.811 | 0.820 | {RESULT} |
47
- | | Sentiment Analysis | 0.777 | 0.781 | 0.790 | {RESULT} |
48
- | **Generation Tasks** | Code Generation | 0.615 | 0.631 | 0.640 | {RESULT} |
49
- | | Creative Writing | 0.588 | 0.579 | 0.601 | {RESULT} |
50
- | | Dialogue Generation | 0.621 | 0.635 | 0.639 | {RESULT} |
51
- | | Summarization | 0.745 | 0.755 | 0.760 | {RESULT} |
52
- | **Specialized Capabilities**| Translation | 0.782 | 0.799 | 0.801 | {RESULT} |
53
- | | Knowledge Retrieval | 0.651 | 0.668 | 0.670 | {RESULT} |
54
- | | Instruction Following | 0.733 | 0.749 | 0.751 | {RESULT} |
55
- | | Safety Evaluation | 0.718 | 0.701 | 0.725 | {RESULT} |
56
 
57
  </div>
58
 
 
38
 
39
  | | Benchmark | Model1 | Model2 | Model1-v2 | MyAwesomeModel |
40
  |---|---|---|---|---|---|
41
+ | **Core Reasoning Tasks** | Math Reasoning | 0.510 | 0.535 | 0.521 | 0.550 |
42
+ | | Logical Reasoning | 0.789 | 0.801 | 0.810 | 0.819 |
43
+ | | Common Sense | 0.716 | 0.702 | 0.725 | 0.736 |
44
+ | **Language Understanding** | Reading Comprehension | 0.671 | 0.685 | 0.690 | 0.700 |
45
+ | | Question Answering | 0.582 | 0.599 | 0.601 | 0.607 |
46
+ | | Text Classification | 0.803 | 0.811 | 0.820 | 0.828 |
47
+ | | Sentiment Analysis | 0.777 | 0.781 | 0.790 | 0.792 |
48
+ | **Generation Tasks** | Code Generation | 0.615 | 0.631 | 0.640 | 0.650 |
49
+ | | Creative Writing | 0.588 | 0.579 | 0.601 | 0.610 |
50
+ | | Dialogue Generation | 0.621 | 0.635 | 0.639 | 0.644 |
51
+ | | Summarization | 0.745 | 0.755 | 0.760 | 0.767 |
52
+ | **Specialized Capabilities**| Translation | 0.782 | 0.799 | 0.801 | 0.804 |
53
+ | | Knowledge Retrieval | 0.651 | 0.668 | 0.670 | 0.676 |
54
+ | | Instruction Following | 0.733 | 0.749 | 0.751 | 0.758 |
55
+ | | Safety Evaluation | 0.718 | 0.701 | 0.725 | 0.739 |
56
 
57
  </div>
58
 
evaluation/utils/__init__.cpython-312-x86_64-linux-gnu.so ADDED
Binary file (23 kB). View file
 
evaluation/utils/benchmark_utils.cpython-312-x86_64-linux-gnu.so ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9a2908e1c33b21a5c30d32106abe1bd1a8237f90766363c3f5ceef24d63bbdee
3
+ size 127208