diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/hf_steered.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/hf_steered.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..c79ddb93e32ba052869ce65ab96eeec1c1b432c1 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/hf_steered.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/openai_completions.cpython-310.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/openai_completions.cpython-310.pyc new file mode 100644 index 0000000000000000000000000000000000000000..4d8699c2fb4512c18de89b32775cefde282cf757 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/openai_completions.cpython-310.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_ipex.cpython-310.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_ipex.cpython-310.pyc new file mode 100644 index 0000000000000000000000000000000000000000..214bbe8c7588ca08824fbb4476a41979eda6f3b4 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_ipex.cpython-310.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_lm.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_lm.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..569ece42e5994565dde42255e2e68cc7be09d869 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/optimum_lm.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/sglang_generate_API.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/sglang_generate_API.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..42b272876083f3e20042a8453161005343152ec5 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/sglang_generate_API.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/textsynth.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/textsynth.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..0254e1b4f5f9166f48c0290fb393ae874f0184aa Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/textsynth.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-310.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-310.pyc new file mode 100644 index 0000000000000000000000000000000000000000..e98584182fb33693d174dc4f39a4ecca14cb53b7 Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-310.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..c3d25a1381a3130fd95b3bb11e409a9e2781f28c Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/utils.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-310.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-310.pyc new file mode 100644 index 0000000000000000000000000000000000000000..4710df95359416799f22cfd92feeca22613ce19d Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-310.pyc differ diff --git a/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-311.pyc b/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-311.pyc new file mode 100644 index 0000000000000000000000000000000000000000..1ca1aa37839b1272c0ae22b6b3f948ad1b5712de Binary files /dev/null and b/lm-evaluation-harness/lm_eval/models/__pycache__/vllm_vlms.cpython-311.pyc differ diff --git a/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/prog.yaml b/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/prog.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c840654f4130e5bb8d47b470445b2318cd8cff44 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/prog.yaml @@ -0,0 +1,12 @@ +task: acp_prog_mcq +dataset_name: acp_prog_mcq +include: _mcq_cot_2shot_yaml +fewshot_config: + sampler: first_n + samples: + - context: 'This is a ferry domain, where the task is to transport cars from their start to their goal locations, using a ferry. Each location is accessible by ferry from each other location. The cars can be debarked or boarded, and the ferry can carry only one car at a time. There are 2 locations and 2 cars, numbered consecutively. Currently, the ferry is at l1, with the car c1 on board. The cars are at locations as follows: c0 is at l1.' + question: 'Which the following facts hold after performing the action \"travel by sea from location l1 to location l0\" in the current state? **Possible Answers**: A. Car c0 is at location l1 and The ferry is at l1 location. B. The ferry is at l0 location and The ferry is at l1 location. C. The ferry is at l0 location. D. The ferry is at l0 location and Car c0 is at location l1.' + answer: "Let's think step by step. Step 1: The following fact(s) do not hold in the current state: The ferry is at l0 location. Step 2: The action adds the following fact(s): The ferry is at l0 location Step 3: The following fact(s) hold in the current state: Car c0 is at location l1. Step 4: The action deletes the following fact(s): The ferry is at l1 location Step 5: Fact(s) \"The ferry is at l0 location\" are added and Fact(s) \"Car c0 is at location l1\" are not deleted. **Final Answer**: D." + - context: 'There are several cities, each containing several locations, some of which are airports. There are also trucks, which can drive within a single city, and airplanes, which can fly between airports. The goal is to get some packages from various locations to various new locations. There are 2 trucks and 1 airplane, as well as 4 packages. There are 4 locations across 2 cities. The locations are in cities as follows: l1-1 and l1-0 are in c1; l0-1 and l0-0 are in c0. Currently, a0 is at l0-0, t1 and p0 are at l1-1, t0 is at l0-1, p1 is in t1, p2 and p3 are in a0.' + question: 'Which the following facts hold after performing the action \"drive truck t0 from location l0-1 in city c0 to location l0-1 in the same city\" in the current state? A. p3 is in t1. B. a0 is at l0-0 and p3 is in t1. C. a0 is at l0-0. D. None of the above.' + answer: "Let's think step by step. Step 1: The following fact(s) hold in the current state: a0 is at l0-0. Step 2: The action deletes the following fact(s): t0 is at l0-1 Step 3: Fact(s) \"a0 is at l0-0\" are not deleted. **Final Answer**: C." diff --git a/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/val.yaml b/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/val.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7aecbc7d5d8e80afab8c05800ef9d507e4cd632a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/acpbench/mcq_cot_2shot/val.yaml @@ -0,0 +1,12 @@ +task: acp_val_mcq +dataset_name: acp_val_mcq +include: _mcq_cot_2shot_yaml +fewshot_config: + sampler: first_n + samples: + - context: 'This is a ferry domain, where the task is to transport cars from their start to their goal locations, using a ferry. Each location is accessible by ferry from each other location. The cars can be debarked or boarded, and the ferry can carry only one car at a time. There are 2 locations and 2 cars, numbered consecutively. Currently, the ferry is at l0 location and it is empty. The cars are at locations as follows: c1 and c0 are at l0. The goal is to reach a state where the following facts hold: Car c0 is at location l1 and Car c1 is at location l1.' + question: 'Which of the following claims is true with regard to the following sequence of actions \"board the car c0 at the location l0, travel by sea from location l0 to location l1, unload the car c0 from the ferry to location l1, travel by sea from location l1 to location l0, board the car c1 at location l0, sail from location l0 to location l1, debark the car c1 from the ferry to location l1\" and the current state? A. The sequence is not applicable. B. The sequence is a plan. C. The sequence is applicable, but does not achieve the goal. D. The sequence is not valid.' + answer: "Let's think step by step. Step 1: For a sequence of actions to be a plan, all actions should be valid, applicable in sequence, and achieve the goal. Step 2: The action sequence is applicable and it achieves the goal. **Final Answer**: B." + - context: 'There are several cities, each containing several locations, some of which are airports. There are also trucks, which can drive within a single city, and airplanes, which can fly between airports. The goal is to get some packages from various locations to various new locations. There are 3 trucks and 1 airplane, as well as 4 packages. There are 9 locations across 3 cities. The locations are in cities as follows: l1-2, l1-0, and l1-1 are in c1; l0-0, l0-1, and l0-2 are in c0; l2-1, l2-2, and l2-0 are in c2. Currently, p2 and t1 are at l1-2, p3 is at l2-0, t0 and p0 are at l0-2, p1 is at l1-0, a0 is at l0-0, t2 is at l2-2. The goal is to reach a state where the following facts hold: p1 is at l1-0, p3 is at l2-0, p2 is at l0-1, and p0 is at l1-2.' + question: 'Which of the following claims is true with regard to the following sequence of actions \"load object p0 into truck t0 at location l0-2, sail the ship t0 into city c0 from location l0-2 in city l0-0, remove the object p0 from the truck t0 and place it on the location l0-0, load the object p0 from location l0-0 onto the airplane a0, fly the airplane a0 from the airport l0-0 to the airport l1-0, remove the object p0 from the airplane a0 and place it on the location l1-0, load object p2 into truck t1 at location l1-2, navigate the truck t1 from its current location l1-2 in city c1 to the new location l1-0 within the same city place the object p0 into the truck t1 at location l1-0 remove the object p2 from the truck t1 and place it on the location l1-0 load the object p2 from location l1-0 onto the airplane a0 fly the airplane a0 from location l1-0 to location l2-0 fly airplane a0 from airport l2-0 to airport l0-0 unload the object p2 from the airplane a0 at location l0-0 place the object p2 into the truck t0 at location l0-0 navigate the truck t0 from its current location l0-0 in city c0 to the new location l0-1 within the same city offload the object p2 from the truck t0 at location l0-1 drive the truck t1 in city c1 from location l1-0 to location l1-2 offload the object p0 from the truck t1 at location l1-2 navigate the truck t2 from its current location l2-2 in city c2 to the new location l2-1 within the same city\" and the current state? A. The sequence is not valid. B. The sequence is applicable, but does not achieve the goal. C. The sequence is a plan. D. The sequence is not applicable.' + answer: "Let's think step by step. Step 1: For a sequence of actions to be a plan, all actions should be valid, applicable in sequence, and achieve the goal. Step 2: The action \"sail the ship t0 into city c0 from location l0-2 in city l0-0\" is not valid in this problem. **Final Answer**: A." diff --git a/lm-evaluation-harness/lm_eval/tasks/aexams/_default_template_yaml b/lm-evaluation-harness/lm_eval/tasks/aexams/_default_template_yaml new file mode 100644 index 0000000000000000000000000000000000000000..3f7100ad70190a67bd86675ce7a15d88a5a5976a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/aexams/_default_template_yaml @@ -0,0 +1,18 @@ +dataset_path: Hennara/aexams +test_split: test +fewshot_split: dev +fewshot_config: + sampler: first_n +output_type: multiple_choice +doc_to_text: "{{question.strip()}}\nA. {{A}}\nB. {{B}}\nC. {{C}}\nD. {{D}}\nالجواب:" +doc_to_choice: ["A", "B", "C", "D"] +doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer)}}" +metric_list: + - metric: acc + aggregation: mean + higher_is_better: true + - metric: acc_norm + aggregation: mean + higher_is_better: true +metadata: + version: 1.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Physics.yaml b/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Physics.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f2764a06ef2680a1c81ccca0e76dcbcf1ba52672 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Physics.yaml @@ -0,0 +1,4 @@ +"dataset_name": "Physics" +"description": "قم بالإجابة على مايلي في مجال الفيزياء \n\n" +"include": "_default_template_yaml" +"task": "aexams_Physics" diff --git a/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Science.yaml b/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Science.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c89dc8c8ca6d32b922483f48ee8da427e027a92b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/aexams/aexams_Science.yaml @@ -0,0 +1,4 @@ +"dataset_name": "Science" +"description": "قم بالإجابة على مايلي في مجال العلوم \n\n" +"include": "_default_template_yaml" +"task": "aexams_Science" diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..bbbb7ef859249191cdf52db9bdab6135319e7a60 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimgsm_yaml +task: afrimgsm_ibo_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..dcfc7160d7b3262ebb78b30a4fa070f748e1e619 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimgsm_yaml +task: afrimgsm_kin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b916281cee9217fafbfdfe62a9c451ef918b36d2 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimgsm_yaml +task: afrimgsm_orm_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..be6dba7151542a8b6ecb4e5cb1da18ab0d9121a3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimgsm_yaml +task: afrimgsm_swa_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..34b773765f90db8695611663bf615744fd6cfaa8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimgsm_yaml +task: afrimgsm_wol_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..d17530bd22284dfd33556f043fb2c18d0325a174 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimgsm_yaml +task: afrimgsm_xho_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_yaml new file mode 100644 index 0000000000000000000000000000000000000000..19d4f7d1fdf39118e2fc774097619b837d34122e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_yaml @@ -0,0 +1,35 @@ +tag: + - afrimgsm_tasks + - afrimgsm_tasks_prompt_1 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +doc_to_text: '{% if answer is not none %}{{question+"\nAnswer:"}}{% else %}{{"Question: "+question+"\nAnswer:"}}{% endif %}' +target_delimiter: "" +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> +filter_list: + - name: remove_whitespace + filter: + - function: remove_whitespace + - function: take_first + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..07b89135ac35c6c3f67df44ff3004e5aeab4197e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_1/afrimgsm_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimgsm_yaml +task: afrimgsm_zul_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_2/afrimgsm_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_2/afrimgsm_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ac0812c1181c50fc51457b65f2cbeb8f64b5a78d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_2/afrimgsm_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimgsm_yaml +task: afrimgsm_amh_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7d214574ade50c17d374e7daad4b4a8e1670db78 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_ibo.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: ibo +doc_to_text: "Answer the given question with the appropriate numerical value, ensuring\ + \ that the response is clear and without any supplementary information. \n\nQuestion:\ + \ {{question}} \nAnswer: " +include: afrimgsm_yaml +task: afrimgsm_ibo_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3d3414e3ca1273567847599c244e266e66e495dc --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_sna.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sna +doc_to_text: "Answer the given question with the appropriate numerical value, ensuring\ + \ that the response is clear and without any supplementary information. \n\nQuestion:\ + \ {{question}} \nAnswer: " +include: afrimgsm_yaml +task: afrimgsm_sna_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..5151f266e673a0106c64c548f53d2b2df18f7896 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_4/afrimgsm_zul.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: zul +doc_to_text: "Answer the given question with the appropriate numerical value, ensuring\ + \ that the response is clear and without any supplementary information. \n\nQuestion:\ + \ {{question}} \nAnswer: " +include: afrimgsm_yaml +task: afrimgsm_zul_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..74131916ab079d9c1bee95d169517661e4d11b31 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lin.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lin +doc_to_text: "For mathematical questions provided in Lingala language. Supply the\ + \ accurate numeric answer to the provided question. \n\nQuestion: {{question}} \n\ + Answer: " +include: afrimgsm_yaml +task: afrimgsm_lin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b92bc4e6574ac7df4aad89c143687747ed5d9863 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_lug.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lug +doc_to_text: "For mathematical questions provided in Luganda language. Supply the\ + \ accurate numeric answer to the provided question. \n\nQuestion: {{question}} \n\ + Answer: " +include: afrimgsm_yaml +task: afrimgsm_lug_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..06cb1b05569674070ee6915f9acf6bd32c6ed72a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_sot.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sot +doc_to_text: "For mathematical questions provided in Sesotho language. Supply the\ + \ accurate numeric answer to the provided question. \n\nQuestion: {{question}} \n\ + Answer: " +include: afrimgsm_yaml +task: afrimgsm_sot_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..0a08c8e3a98efcc19a3f48ec4dc46d220fd7a9ab --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_swa.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: swa +doc_to_text: "For mathematical questions provided in Swahili language. Supply the\ + \ accurate numeric answer to the provided question. \n\nQuestion: {{question}} \n\ + Answer: " +include: afrimgsm_yaml +task: afrimgsm_swa_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_vai.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_vai.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ab3337a7eaf4f7f3a4d06945fd56b78b7d74763f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_vai.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: vai +doc_to_text: "For mathematical questions provided in Vai language. Supply the accurate\ + \ numeric answer to the provided question. \n\nQuestion: {{question}} \nAnswer: " +include: afrimgsm_yaml +task: afrimgsm_vai_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cd0bea64173d14db8bc6d472bb685ee9f7b420a8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct/prompt_5/afrimgsm_yor.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: yor +doc_to_text: "For mathematical questions provided in Yoruba language. Supply the accurate\ + \ numeric answer to the provided question. \n\nQuestion: {{question}} \nAnswer: " +include: afrimgsm_yaml +task: afrimgsm_yor_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..57c0e564b1b594fe21203a0d73737301167b63cb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_eng.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: eng +include: afrimgsm_cot_yaml +task: afrimgsm_cot_eng_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..55fdff7c365f895a4835b6234a7027a28b7c42a3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ewe_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f42e0ee559e562a934c3151cc9fd535442c6cf28 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimgsm_cot_yaml +task: afrimgsm_cot_hau_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..dfabc3e319236dd2112ab74bbb5d1181f63d8d55 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ibo_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..24f553497c41bb4c42b99d9ba33abc21d42ebf15 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimgsm_cot_yaml +task: afrimgsm_cot_twi_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_vai.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_vai.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cc63717012dd2df5a3965b352363ac9a2457372f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_vai.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: vai +include: afrimgsm_cot_yaml +task: afrimgsm_cot_vai_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c86b09d6a9961af49a771afd9ef62f991edc8084 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimgsm_cot_yaml +task: afrimgsm_cot_wol_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6f03080c3c6c986c760578e766e8d23c2e58a17f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimgsm_cot_yaml +task: afrimgsm_cot_xho_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_yaml new file mode 100644 index 0000000000000000000000000000000000000000..6ab733bf114ba32013ab433ca74a1f05b66f8a78 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_yaml @@ -0,0 +1,37 @@ +tag: + - afrimgsm_cot_tasks + - afrimgsm_cot_tasks_prompt_1 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +training_split: train +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +doc_to_text: '{% if answer is not none %}{{question+"\nStep-by-Step Answer:"}}{% else %}{{"Question: "+question+"\nStep-by-Step Answer:"}}{% endif %}' +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> + - <|eot_id|> +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +filter_list: + - name: "strict-match" + filter: + - function: "regex" + regex_pattern: "The answer is (\\-?[0-9\\.\\,]+)" + - function: "take_first" + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..5cacc2be5c426dde5bdfeb5dc5e3e62b20e494b8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_1/afrimgsm_cot_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimgsm_cot_yaml +task: afrimgsm_cot_zul_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6d5d43fb8562da74b469cc15636aaac35c082ef2 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimgsm_cot_yaml +task: afrimgsm_cot_amh_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..84a6b26dd7d5f17dc76f2bc5c28bfa05d0936805 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_eng.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: eng +include: afrimgsm_cot_yaml +task: afrimgsm_cot_eng_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..987ac630ee57139c93588b86dc8ae53bbd6c466e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimgsm_cot_yaml +task: afrimgsm_cot_fra_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..488f693a60c4a31cec8adf1e2665cf7cd6430b18 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimgsm_cot_yaml +task: afrimgsm_cot_hau_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..aefa0aa229981bd07b9e221a66d4ebce08c0f086 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ibo_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e183dcd48f80b358ed54bdde2fd8d9e3ac64a9ba --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimgsm_cot_yaml +task: afrimgsm_cot_kin_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a36d89355bc25c75f7097c371ed685c1f05a8290 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimgsm_cot_yaml +task: afrimgsm_cot_orm_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..25187ceccadcbb09e0a8bf417d5b2ae4b4572b0d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sna_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b0d91d0c9e3cebf8352c589c4ab6ac7a30183129 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimgsm_cot_yaml +task: afrimgsm_cot_swa_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..03b59a394366bca2b5241bd3a096b7b274354cad --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimgsm_cot_yaml +task: afrimgsm_cot_twi_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_vai.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_vai.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8fa4cf5e2db3121bae1a296472cefecc34c31ceb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_vai.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: vai +include: afrimgsm_cot_yaml +task: afrimgsm_cot_vai_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2611de84f42fd9a0804cb155f94b687bab875a81 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimgsm_cot_yaml +task: afrimgsm_cot_wol_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..33059776d4a77e62c6a0478efa64ecd8da5c6a0f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimgsm_cot_yaml +task: afrimgsm_cot_xho_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yaml new file mode 100644 index 0000000000000000000000000000000000000000..505336ba01a57df47f10104bedb8af7288e4d98d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yaml @@ -0,0 +1,37 @@ +tag: + - afrimgsm_cot_tasks + - afrimgsm_cot_tasks_prompt_2 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +training_split: train +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +doc_to_text: 'Give direct numerical answers for the question provided. \n\nQuestion: {{question}} \Step-by-Step Answer: ' +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> + - <|eot_id|> +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +filter_list: + - name: "strict-match" + filter: + - function: "regex" + regex_pattern: "The answer is (\\-?[0-9\\.\\,]+)" + - function: "take_first" + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..991297c4dd40aaeab196f45fbf6dcb6521432c25 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_2/afrimgsm_cot_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimgsm_cot_yaml +task: afrimgsm_cot_yor_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..00f830a20eb0e580600868cd88ca9d25231352c1 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimgsm_cot_yaml +task: afrimgsm_cot_amh_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ea0937f233d75fea06dde81274381f2085a4416e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_eng.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: eng +include: afrimgsm_cot_yaml +task: afrimgsm_cot_eng_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..dfe111d7fc2b3ebc91463719305e734c02a2360f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ewe_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..eb82d3a44416adc5abb640d460d49c428e71f1bf --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimgsm_cot_yaml +task: afrimgsm_cot_fra_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3162114b1eeb9a931ad6c006636fa8d57605c272 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimgsm_cot_yaml +task: afrimgsm_cot_hau_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f46191a331e0498b6a93cdda5f51750b42702423 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ibo_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8ddc82ee85066be8890b863bc042bc66cc44316e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimgsm_cot_yaml +task: afrimgsm_cot_kin_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..769ae73aa82813832af62ccf69a1e37e0bf688ae --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimgsm_cot_yaml +task: afrimgsm_cot_lin_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e04769a6c77dc035b03e48dfffb6fc2ef90081b6 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimgsm_cot_yaml +task: afrimgsm_cot_lug_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..79a696581beac371a7afbf71e918d6e6f43263c5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimgsm_cot_yaml +task: afrimgsm_cot_orm_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f08d44259f104d9b72c5ed732c0dbb6938524703 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sna_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..76ea5f96ae6a2618b54a1009684d8e632c9ebd0e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimgsm_cot_yaml +task: afrimgsm_cot_swa_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c45b3f0fcfab17390f077b2d161f76f097dfda33 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimgsm_cot_yaml +task: afrimgsm_cot_twi_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_vai.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_vai.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ca50c481fd6c8470bc6d61e012969033cbf2bcfb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_vai.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: vai +include: afrimgsm_cot_yaml +task: afrimgsm_cot_vai_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..16dbc506ef36a0799834d1cd5af1166d029453d3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimgsm_cot_yaml +task: afrimgsm_cot_wol_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a329b8ebf4487eb5b922f539274b6592d6d2e75f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimgsm_cot_yaml +task: afrimgsm_cot_xho_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yaml new file mode 100644 index 0000000000000000000000000000000000000000..d4d3657da5f68eb670173ff86034aa2276c2c0ae --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yaml @@ -0,0 +1,37 @@ +tag: + - afrimgsm_cot_tasks + - afrimgsm_cot_tasks_prompt_3 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +training_split: train +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +doc_to_text: 'Solve the following math question \n\nQuestion: {{question}} \nStep-by-Step Answer: ' +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> + - <|eot_id|> +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +filter_list: + - name: "strict-match" + filter: + - function: "regex" + regex_pattern: "The answer is (\\-?[0-9\\.\\,]+)" + - function: "take_first" + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..003fb63482132f44de4252d96afb25808f99c876 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimgsm_cot_yaml +task: afrimgsm_cot_yor_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c01468ec7f6a5e6d87f132599fc5df4879516c95 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_3/afrimgsm_cot_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimgsm_cot_yaml +task: afrimgsm_cot_zul_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6624ddfe5319d574caf011f2de342bb68d516771 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_amh.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: amh +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_amh_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..135bd975b0a9a4a37f2d3fcce07c6d96b9c9579e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ewe.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: ewe +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ewe_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..81a060b2c2b7ce84836f88aac2e5e386c2ad2e6b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_fra.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: fra +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_fra_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b53dba5852f274ada14954ef6b839f288d1629dd --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_hau.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: hau +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_hau_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2a4236e1d6a36e7d241c9d65ca1c111b8bdac536 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_ibo.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: ibo +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ibo_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..51407a6626539a7a79899f5691122c4c5c0881db --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_kin.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: kin +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_kin_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..248ffeee0130f295668f8b9596dea68b6b077527 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lin.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lin +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_lin_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..fbf7c8cd5edf41c05c347423dcb61cb5f84420f3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_lug.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lug +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_lug_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..218c3f90a1882f702fed06585f30694ef8e9e96b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_orm.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: orm +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_orm_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..81e4840a177d442dd62d2fd08a3e3b36458a65b0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sna.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sna +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sna_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..47bcd414523bf5a4b81440e08476c9f9a2d4e794 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_sot.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sot +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sot_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e0b57a14edab9695ceda40dc5a1d44fcd0eb2230 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_swa.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: swa +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_swa_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..abdbdec70ee3203fd39cc99293c52aecc88f37e0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_twi.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: twi +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_twi_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_vai.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_vai.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a0b7913b381f35893889225d9cfd3aaaf25555c9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_vai.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: vai +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_vai_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..aa75a3f599284d5dfc9dd58f4647140180e93378 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_wol.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: wol +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_wol_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8c125ebedfab2aa922c27627ec29b582b3d6fa37 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_xho.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: xho +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_xho_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yaml new file mode 100644 index 0000000000000000000000000000000000000000..59013d84ed916ab9728f3345f6323b9fbee4c8d6 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yaml @@ -0,0 +1,36 @@ +tag: + - afrimgsm_cot_tasks + - afrimgsm_cot_tasks_prompt_4 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +training_split: train +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> + - <|eot_id|> +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +filter_list: + - name: "strict-match" + filter: + - function: "regex" + regex_pattern: "The answer is (\\-?[0-9\\.\\,]+)" + - function: "take_first" + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2c960b75f60dcfef53fac9f5cff333baeeb00ef2 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_4/afrimgsm_cot_yor.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: yor +doc_to_text: "Answer the given question with the step by step solution appropriate\ + \ numerical value, ensuring that the response is clear and without any supplementary\ + \ information. \n\nQuestion: {{question}} \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_yor_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ea5124850e7ce8cdf270c5cd53020c61d9e10491 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_amh.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: amh +doc_to_text: "For mathematical questions provided in Amharic language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_amh_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..9b485061e5974c691035eaaf18af67001b921e73 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_eng.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: eng +doc_to_text: "For mathematical questions provided in English language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_eng_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e52f43276704281a95bd366f8512ca6c9af3f4c3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_ewe.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: ewe +doc_to_text: "For mathematical questions provided in Ewe language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_ewe_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f311e12a3b8bbb77e9f24d11fa15efbea38330f5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_fra.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: fra +doc_to_text: "For mathematical questions provided in French language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_fra_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..91cc7ace3922ace059d7c4f789d70ba1162d183b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_hau.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: hau +doc_to_text: "For mathematical questions provided in Hausa language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_hau_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..419da8ab7aebe31e28c40bb9bd80b34e5b87f867 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_lin.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lin +doc_to_text: "For mathematical questions provided in Lingala language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_lin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a9a448f2571248957e397d2660d6bc019ac9245d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_orm.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: orm +doc_to_text: "For mathematical questions provided in Oromo language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_orm_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..645b2898c4ad5ec017af12c813ea3d27f1231d55 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sna.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sna +doc_to_text: "For mathematical questions provided in chiShona language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sna_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a0b940d919edb0a710784ffcbdab00110b568a65 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_sot.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sot +doc_to_text: "For mathematical questions provided in Sesotho language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_sot_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..093ccfa2f705e7f94cffed63e0406f0653f1c867 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_swa.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: swa +doc_to_text: "For mathematical questions provided in Swahili language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_swa_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b73863adafc318c81c058efdfb3fd251cb987a95 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_wol.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: wol +doc_to_text: "For mathematical questions provided in Wolof language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_wol_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1b77d56f2115fc188b5d973d5aaeb33b506830b5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_xho.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: xho +doc_to_text: "For mathematical questions provided in isiXhosa language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_xho_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yaml new file mode 100644 index 0000000000000000000000000000000000000000..de15089149d133ac1f84ee89d1a20634286ed10c --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yaml @@ -0,0 +1,36 @@ +tag: + - afrimgsm_cot_tasks + - afrimgsm_cot_tasks_prompt_5 +dataset_path: masakhane/afrimgsm +dataset_name: null # Overridden by language-specific config. +output_type: generate_until +training_split: train +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> + - <|eot_id|> +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +filter_list: + - name: "strict-match" + filter: + - function: "regex" + regex_pattern: "The answer is (\\-?[0-9\\.\\,]+)" + - function: "take_first" + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..9032313ad485bd0d9e2b90854bb626b432dc1a46 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/direct_cot/prompt_5/afrimgsm_cot_yor.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: yor +doc_to_text: "For mathematical questions provided in Yoruba language. Supply the accurate\ + \ step by step answer to the provided question. \n\nQuestion: {{question}} \nStep\ + \ by step answer: " +include: afrimgsm_cot_yaml +task: afrimgsm_cot_yor_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/gen_yaml.sh b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/gen_yaml.sh new file mode 100644 index 0000000000000000000000000000000000000000..5c0132822a7f3ba68230762e0342838583c29bd9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/gen_yaml.sh @@ -0,0 +1,7 @@ +#!/bin/bash + +# python utils.py --overwrite --output-dir direct --mode direct +# python utils.py --overwrite --output-dir direct_native --mode direct-native +# python utils.py --overwrite --output-dir en_cot --mode en-cot +# python utils.py --overwrite --output-dir native_cot --mode native-cot +python utils.py --overwrite --output-dir translate_direct --mode translate-direct diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/run.sh b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/run.sh new file mode 100644 index 0000000000000000000000000000000000000000..075500be33775dc49288ce7f7180604c7c6f99ce --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/run.sh @@ -0,0 +1,6 @@ +lm_eval --model hf \ + --model_args pretrained="google/gemma-7b" --tasks afrimgsm_en_cot_eng,mgsm_en_cot_en,afrimgsm_native_cot_eng,mgsm_native_cot_en,afrimgsm_direct_eng,mgsm_direct_en,afrimgsm_direct_native_eng \ + --device cuda:0 \ + --batch_size 1 \ + --verbosity DEBUG \ + --limit 5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/afrimgsm_tt.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/afrimgsm_tt.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e1cc68abefd07d38ef64c3524337be287a20e779 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/afrimgsm_tt.yaml @@ -0,0 +1,9 @@ +group: afrimgsm_tt-irokobench +task: + - afrimgsm_tt_tasks +aggregate_metric_list: + - metric: acc + aggregation: mean + weight_by_size: true +metadata: + version: 2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..b38e82f252e4379f8332ba7b779617db5d53da24 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimgsm_translate_yaml +task: afrimgsm_translate_fra_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..5333b163698e15aa1b2548395dfa1069e188b6b1 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimgsm_translate_yaml +task: afrimgsm_translate_ibo_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ae231d6da0ca0de8571eaaf38cb62121fe57d125 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimgsm_translate_yaml +task: afrimgsm_translate_kin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..65349c7e66fbda73b3262d21fba18c69a88a318a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimgsm_translate_yaml +task: afrimgsm_translate_lin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7643fc1223f63f698a8ef70835beccea7de9d7a5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimgsm_translate_yaml +task: afrimgsm_translate_lug_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..55e1992799923fc30ffea628a64b1d4bcab31bf0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimgsm_translate_yaml +task: afrimgsm_translate_orm_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2b206e3fc82f9e519aece099aa2e5d31c783df62 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_sot.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sot +include: afrimgsm_translate_yaml +task: afrimgsm_translate_sot_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1abdd50bced23bc6de7768b9f2db931c7f35b4ad --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimgsm_translate_yaml +task: afrimgsm_translate_xho_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..d3927ba8e0369a5220df6d72e4bb474b7e8af7ca --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimgsm_translate_yaml +task: afrimgsm_translate_yor_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a57260d94ac8b83f7371baecab3e822868dc671b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_1/afrimgsm_translate_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimgsm_translate_yaml +task: afrimgsm_translate_zul_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..86d8dbbca635d347640c699db7bde0cb6d950722 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimgsm_translate_yaml +task: afrimgsm_translate_hau_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..53078341b5822305b3cc5de3652f0e582313aff6 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimgsm_translate_yaml +task: afrimgsm_translate_kin_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..137ccbcd322421d17c4ea369c34f1beac05ce597 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimgsm_translate_yaml +task: afrimgsm_translate_sna_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_yaml new file mode 100644 index 0000000000000000000000000000000000000000..63766339e6e0cb5bbb3d9f45d1010d00de0aafd4 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate/prompt_2/afrimgsm_translate_yaml @@ -0,0 +1,34 @@ +tag: afrimgsm_tt_tasks +dataset_path: masakhane/afrimgsm-translate-test +output_type: generate_until +test_split: test +doc_to_target: '{% if answer is not none %}{{answer[21:]}}{% else %}{{answer_number|string}}{% endif %}' +doc_to_text: "Give direct numerical answers for the question provided. \n\nQuestion: {{question}} \nAnswer: " +target_delimiter: "" +generation_kwargs: + do_sample: false + until: + - 'Question:' + - + - <|im_end|> +should_decontaminate: true +doc_to_decontamination_query: "Answer: " +filter_list: + - name: remove_whitespace + filter: + - function: remove_whitespace + - function: take_first + - filter: + - function: regex + group_select: -1 + regex_pattern: (-?[$0-9.,]{2,})|(-?[0-9]+) + - function: take_first + name: flexible-extract +metric_list: + - metric: exact_match + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true +metadata: + version: 2.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..87db46c9761e033a42739d1aaaf1d14e51989d14 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_lin.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: lin +doc_to_text: "For mathematical questions provided in Lingala language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_translate_yaml +task: afrimgsm_cot_translate_lin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..d5eac98d4e45e5c199cc5a19286c7d0edcd04a9f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimgsm/translate_cot/prompt_5/afrimgsm_cot_translate_sna.yaml @@ -0,0 +1,7 @@ +# Generated by utils.py +dataset_name: sna +doc_to_text: "For mathematical questions provided in chiShona language. Supply the\ + \ accurate step by step answer to the provided question. \n\nQuestion: {{question}}\ + \ \nStep by step answer: " +include: afrimgsm_cot_translate_yaml +task: afrimgsm_cot_translate_sna_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/afrimmlu.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/afrimmlu.yaml new file mode 100644 index 0000000000000000000000000000000000000000..202c31825bfcdaa8ea974e8f51444bc864ed4306 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/afrimmlu.yaml @@ -0,0 +1,13 @@ +group: afrimmlu-irokobench +task: + - afrimmlu_tasks_prompt_1 + - afrimmlu_tasks_prompt_2 + - afrimmlu_tasks_prompt_3 + - afrimmlu_tasks_prompt_4 + - afrimmlu_tasks_prompt_5 +aggregate_metric_list: + - metric: acc + aggregation: mean + weight_by_size: true +metadata: + version: 2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8a26369b36ee47ed6ac21c448c316acaf90af749 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_direct +task: afrimmlu_direct_amh_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..4e8a2875e71f43dbdd148331d24e6440f92ad71f --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimmlu_direct +task: afrimmlu_direct_fra_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8b08d48e0a3c42e123c081935d5ffcc71e1c56c7 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimmlu_direct +task: afrimmlu_direct_ibo_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..00d82dfa57e6a77d52f476ae78c54edbc677628d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_direct +task: afrimmlu_direct_kin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3301647298d216de664ff07e2c8a10e134afe388 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimmlu_direct +task: afrimmlu_direct_lug_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c5ebed9f96e857356adf2ccaa2de2cf818874e71 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_direct +task: afrimmlu_direct_swa_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..4ccbc47cd02c5de0c886deed9f4549141884eac8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimmlu_direct +task: afrimmlu_direct_wol_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3e30d2017740585b5c675ec898fa9d9512e4ac52 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/afrimmlu_direct_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimmlu_direct +task: afrimmlu_direct_xho_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..f1bb9162f0fbc68807db68134970ae2636980cbf --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_1/utils.py @@ -0,0 +1,32 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_choice(doc): + choices = eval(doc["choices"]) + return choices + + +def doc_to_text(doc): + output = """You are a highly knowledgeable and intelligent artificial intelligence + model answers multiple-choice questions about {subject} + + Question: {question} + + Choices: + A: {choice1} + B: {choice2} + C: {choice3} + D: {choice4} + + Answer: """ + + choices = eval(doc["choices"]) + text = output.format( + subject=doc["subject"], + question=doc["question"], + choice1=choices[0], + choice2=choices[1], + choice3=choices[2], + choice4=choices[3], + ) + return text diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..85d85171bfe9d84205a9ab218ed496aed1eecf73 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_direct +task: afrimmlu_direct_amh_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..47f0bfb14b6cd7eaf34618eb0709f9f3f0c9b666 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimmlu_direct +task: afrimmlu_direct_fra_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..29b4a4d2029e945c4bf52654a819dcb89b898431 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimmlu_direct +task: afrimmlu_direct_hau_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ce7c2e896509b874459deb156f2ba34287a908c9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_direct +task: afrimmlu_direct_kin_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..51fcea62af8ca4fa64643ef4ea170444ce25beef --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_direct +task: afrimmlu_direct_lin_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f4c57ae36fbad9a75737562b6cb619b53e933c34 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimmlu_direct +task: afrimmlu_direct_lug_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..494d4240693fe9907f965cd8ad5ccc71fcfc2868 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimmlu_direct +task: afrimmlu_direct_orm_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7706ad64ccc0baef8fc4964e61871805187db548 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimmlu_direct +task: afrimmlu_direct_sna_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..353bd2574657f0b7f49d0be77edd777893ac549b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_sot.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sot +include: afrimmlu_direct +task: afrimmlu_direct_sot_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..54a16c6c2f5839c6aec545c74f6a5df3293938df --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_direct +task: afrimmlu_direct_swa_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8bb35bd5f9ad784fbf2101e3ce14e82710c75858 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimmlu_direct +task: afrimmlu_direct_twi_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8766392a0dc7e8abd5b8145f37bcd13f424a8b6a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/afrimmlu_direct_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_direct +task: afrimmlu_direct_zul_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..e0cfb334c27cfe4c5bbb1ff7126215c0ea9130c9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_2/utils.py @@ -0,0 +1,30 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_choice(doc): + choices = eval(doc["choices"]) + return choices + + +def doc_to_text(doc): + output = """As an expert in {subject}, choose the most accurate answer to the question below. +Your goal is to select the correct option 'A', 'B', 'C', or 'D' by understanding the nuances of the topic. + +Question: {question} +Choices: + A: {choice1} + B: {choice2} + C: {choice3} + D: {choice4} +Answer: """ + + choices = eval(doc["choices"]) + text = output.format( + subject=doc["subject"], + question=doc["question"], + choice1=choices[0], + choice2=choices[1], + choice3=choices[2], + choice4=choices[3], + ) + return text diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct new file mode 100644 index 0000000000000000000000000000000000000000..fb2fd165fcba0457c82e825afa5d8252546dc09c --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct @@ -0,0 +1,37 @@ +tag: + - afrimmlu_tasks + - afrimmlu_tasks_prompt_3 + - afrobench_mmlu_tasks +dataset_path: masakhane/afrimmlu +dataset_name: null +output_type: multiple_choice +validation_split: validation +test_split: test +fewshot_split: validation +doc_to_text: !function utils.doc_to_text +doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer)}}" +doc_to_choice: !function utils.doc_to_choice +should_decontaminate: true +doc_to_decontamination_query: "Question: {{question}}\nAnswer:" +metric_list: + - metric: f1 + aggregation: !function utils.weighted_f1_score + # aggregation: mean + average: weighted + hf_evaluate: true + higher_is_better: True + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" + - metric: acc + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" +metadata: + version: 1.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c7c28f20b09193f8a0a5c1c0f4ffd8ae59312a08 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_direct +task: afrimmlu_direct_amh_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..83f7cfcb32c1d85061a3d9b6e1cca169a61d4ff0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_eng.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: eng +include: afrimmlu_direct +task: afrimmlu_direct_eng_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..351bdf330c4b30d85448e45b9233aaf6cb704c4b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimmlu_direct +task: afrimmlu_direct_ewe_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..691978187578805d95bc215c8d678273f02d343d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimmlu_direct +task: afrimmlu_direct_fra_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..90521523bef4afe8420fc579108dd2353afb49f3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimmlu_direct +task: afrimmlu_direct_hau_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..43a88fe6c13d804531aa7e251fe86c0102562bc9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimmlu_direct +task: afrimmlu_direct_ibo_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..977f3ab259efac5e6afcccdfd44e04279651ad18 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_direct +task: afrimmlu_direct_kin_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2d25584a3fe0d9c7a73e13ab7a3f1b8652616efd --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_direct +task: afrimmlu_direct_lin_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2b4da1a7f550f66c1b3f084879413d2d9fc13641 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimmlu_direct +task: afrimmlu_direct_lug_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2738f980d41cea65fa790aa16232a0f6a7584226 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimmlu_direct +task: afrimmlu_direct_orm_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..063d111ac4aa0c8625d8615e7b13c1d10ac906fb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimmlu_direct +task: afrimmlu_direct_sna_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6cf6e66d4fb91e3e98ed5b0fe955379bcb31bf26 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_sot.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sot +include: afrimmlu_direct +task: afrimmlu_direct_sot_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e90204d40f2fef22b46b190b552cd9f9fcb777b0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_direct +task: afrimmlu_direct_swa_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..719ebe9002cc9fab19dfa793153d779ad8ffbee6 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimmlu_direct +task: afrimmlu_direct_twi_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8f0f1d0d709b8f0dddd5708b66f6d27a122984d0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimmlu_direct +task: afrimmlu_direct_wol_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8fc1af4d171021226243c69a59b44363b4a16639 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimmlu_direct +task: afrimmlu_direct_xho_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..a641b03ae173f2342e8d7178119e68fea2e5f000 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimmlu_direct +task: afrimmlu_direct_yor_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8c6b493d34a99a4250676c2f0130ceff6b4ea4f8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/afrimmlu_direct_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_direct +task: afrimmlu_direct_zul_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..bc3da2e29667b4b25f68757e2169a5c8aa0c8dea --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_3/utils.py @@ -0,0 +1,32 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_choice(doc): + choices = eval(doc["choices"]) + return choices + + +def doc_to_text(doc): + output = """You are a subject matter expert in {subject}. + + Utilizing your expertise in {subject}, answer the following multiple-choice question + by picking 'A', 'B', 'C', or 'D'. + +Question: {question} +Choices: + A: {choice1} + B: {choice2} + C: {choice3} + D: {choice4} +Answer: """ + + choices = eval(doc["choices"]) + text = output.format( + subject=doc["subject"], + question=doc["question"], + choice1=choices[0], + choice2=choices[1], + choice3=choices[2], + choice4=choices[3], + ) + return text diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct new file mode 100644 index 0000000000000000000000000000000000000000..c15b7b2fc3991517b15f2c370a246adb907f2e52 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct @@ -0,0 +1,37 @@ +tag: + - afrimmlu_tasks + - afrimmlu_tasks_prompt_4 + - afrobench_mmlu_tasks +dataset_path: masakhane/afrimmlu +dataset_name: null +output_type: multiple_choice +validation_split: validation +test_split: test +fewshot_split: validation +doc_to_text: !function utils.doc_to_text +doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer)}}" +doc_to_choice: !function utils.doc_to_choice +should_decontaminate: true +doc_to_decontamination_query: "Question: {{question}}\nAnswer:" +metric_list: + - metric: f1 + aggregation: !function utils.weighted_f1_score + # aggregation: mean + average: weighted + hf_evaluate: true + higher_is_better: True + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" + - metric: acc + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" +metadata: + version: 1.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cc862dc2327d23f394742ee51031e37fc7c97ff1 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_direct +task: afrimmlu_direct_amh_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..69baef502b78bce307b77a7c44c8c4323ebc1102 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_eng.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: eng +include: afrimmlu_direct +task: afrimmlu_direct_eng_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f5af1074f4b993daa4f2468f60baf83b48bfc470 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimmlu_direct +task: afrimmlu_direct_ewe_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ca8f7c5ed0ae15bc5a5e96c776f2251f2cff06fa --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimmlu_direct +task: afrimmlu_direct_hau_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8a181d07cc6a81aaf42308fc137c2241a0d8d444 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimmlu_direct +task: afrimmlu_direct_ibo_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8f86122466a7db1fdc06a2352385fdc1fc78bd69 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_direct +task: afrimmlu_direct_kin_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3c7d3ecf7a86a68248a49d5fc97947ca8da69b0b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_direct +task: afrimmlu_direct_lin_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..467201319f12257549de2fa3c260591dee13f311 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimmlu_direct +task: afrimmlu_direct_lug_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e52668253d495c2583d9f5e964dc73c6850a5729 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimmlu_direct +task: afrimmlu_direct_orm_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..af29225a1ca8c63c042d661ce5dc5331a44ed28a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimmlu_direct +task: afrimmlu_direct_sna_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..0342dc10b71f69880b3a3352f9d9def10f9815c8 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_sot.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sot +include: afrimmlu_direct +task: afrimmlu_direct_sot_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ec9a3525f534ac5ea7ed6817d8f52bc67b57c445 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_direct +task: afrimmlu_direct_swa_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..83dc916c68120a6a28234dfaad48b71c9cbfdba3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimmlu_direct +task: afrimmlu_direct_twi_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e656af2c6697004f7ea94ff638b8b8d4f9f8d549 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimmlu_direct +task: afrimmlu_direct_wol_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ab23d9346400d5ec0eedb7f91f530a04499f163c --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimmlu_direct +task: afrimmlu_direct_xho_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..0dd0254819a9ef2b6b2b795bbcfad7ed8ef1c314 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimmlu_direct +task: afrimmlu_direct_yor_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..98a0937fc74a67a0fae7faa39164f10c288aa3f7 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/afrimmlu_direct_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_direct +task: afrimmlu_direct_zul_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..29c23b7f856b2ab4ead359cafbbf404241e53ffb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_4/utils.py @@ -0,0 +1,28 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_choice(doc): + choices = eval(doc["choices"]) + return choices + + +def doc_to_text(doc): + output = """Analyze each question critically and determine the most correct option based on your understanding of the subject matter + +Question: {question} +Choices: + A: {choice1} + B: {choice2} + C: {choice3} + D: {choice4} +Answer: """ + + choices = eval(doc["choices"]) + text = output.format( + question=doc["question"], + choice1=choices[0], + choice2=choices[1], + choice3=choices[2], + choice4=choices[3], + ) + return text diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct new file mode 100644 index 0000000000000000000000000000000000000000..3da1eb827af65c9bcb69dd4af7eab06df848ade2 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct @@ -0,0 +1,37 @@ +tag: + - afrimmlu_tasks + - afrimmlu_tasks_prompt_5 + - afrobench_mmlu_tasks +dataset_path: masakhane/afrimmlu +dataset_name: null +output_type: multiple_choice +validation_split: validation +test_split: test +fewshot_split: validation +doc_to_text: !function utils.doc_to_text +doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer)}}" +doc_to_choice: !function utils.doc_to_choice +should_decontaminate: true +doc_to_decontamination_query: "Question: {{question}}\nAnswer:" +metric_list: + - metric: f1 + aggregation: !function utils.weighted_f1_score + # aggregation: mean + average: weighted + hf_evaluate: true + higher_is_better: True + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" + - metric: acc + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" +metadata: + version: 1.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cff031d7936a5b82bacf175d8d17e54a51d7fe92 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_direct +task: afrimmlu_direct_amh_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cef2f86599b589c38e7fb30723a3024b3fcfeffe --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimmlu_direct +task: afrimmlu_direct_ewe_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..042c0bbbf0c5afc60776127770b88d15c7c7c5ea --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrimmlu_direct +task: afrimmlu_direct_fra_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_hau.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_hau.yaml new file mode 100644 index 0000000000000000000000000000000000000000..cd507182558a8884859713d4fbcf356898d9176c --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_hau.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: hau +include: afrimmlu_direct +task: afrimmlu_direct_hau_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..9e9839001ed95c56ae13b1fa97466fbbb39c4acb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimmlu_direct +task: afrimmlu_direct_ibo_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1d157559f8c8ea7172bdf64a78506e02830ee633 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_direct +task: afrimmlu_direct_kin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6eca1f8e7ce1b442be6b62ec924143d736f97c62 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_direct +task: afrimmlu_direct_lin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..854b160dc4dc6bc4192ab6c2f73a0e8286da6376 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_lug.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lug +include: afrimmlu_direct +task: afrimmlu_direct_lug_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..9592e585bbe1c62b6b4f2eb25ab02facda7bc242 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimmlu_direct +task: afrimmlu_direct_orm_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..51d05c686db303d55507c16f8a31c56ec3222f29 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimmlu_direct +task: afrimmlu_direct_sna_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c1cd2672b09792c6fd0cf6a9c3c77b83ac5cdbcb --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_direct +task: afrimmlu_direct_swa_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1e2e258c6d020e6599fe3d8fd92388958fce14b1 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimmlu_direct +task: afrimmlu_direct_twi_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..d721871b35a08b564c04af09ca32349a4433bf93 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrimmlu_direct +task: afrimmlu_direct_wol_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..8f528abb1018753945b954945138084b2d7327ce --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimmlu_direct +task: afrimmlu_direct_yor_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ec83abebdb65535e345a1c488c2e2999f798d373 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/afrimmlu_direct_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_direct +task: afrimmlu_direct_zul_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..a47ceca967c136c7df7132d826ac51af26039722 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/direct/prompt_5/utils.py @@ -0,0 +1,29 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_choice(doc): + choices = eval(doc["choices"]) + return choices + + +def doc_to_text(doc): + output = """Given your proficiency in {subject}, please answer the subsequent multiple-choice question with 'A', 'B', 'C', or 'D'. + +Question: {question} +Choices: + A: {choice1} + B: {choice2} + C: {choice3} + D: {choice4} +Answer: """ + + choices = eval(doc["choices"]) + text = output.format( + subject=doc["subject"], + question=doc["question"], + choice1=choices[0], + choice2=choices[1], + choice3=choices[2], + choice4=choices[3], + ) + return text diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/gen_utils.py b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/gen_utils.py new file mode 100644 index 0000000000000000000000000000000000000000..a195b6b5852d35042c14632597762a3965faae07 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/gen_utils.py @@ -0,0 +1,103 @@ +import argparse +import os + +import yaml + + +class FunctionTag: + def __init__(self, value): + self.value = value + + +def gen_lang_yamls(output_dir: str, overwrite: bool, mode: str) -> None: + """ + Generate a yaml file for each language. + + :param output_dir: The directory to output the files to. + :param overwrite: Whether to overwrite files if they already exist. + """ + err = [] + languages = { + "eng": "English", + "amh": "Amharic", + "ibo": "Igbo", + "fra": "French", + "sna": "chiShona", + "wol": "Wolof", + "ewe": "Ewe", + "lin": "Lingala", + "lug": "Luganda", + "xho": "isiXhosa", + "kin": "Kinyarwanda", + "twi": "Twi", + "zul": "Zulu", + "orm": "Oromo", + "yor": "Yoruba", + "hau": "Hausa", + "sot": "Sesotho", + "swa": "Swahili", + } + + for lang in languages.keys(): + try: + file_name = f"afrimmlu_direct_{lang}.yaml" + task_name = f"afrimmlu_direct_{lang}_{mode}" + yaml_template = "afrimmlu_direct" + if output_dir.split("/")[-1] == "translate": + file_name = f"afrimmlu_translate_{lang}.yaml" + task_name = f"afrimmlu_translate_{lang}_{mode}" + yaml_template = "afrimmlu_translate" + yaml_details = { + "include": yaml_template, + "task": task_name, + "dataset_name": lang, + } + os.makedirs(f"{output_dir}/{mode}", exist_ok=True) + with open( + f"{output_dir}/{mode}/{file_name}", + "w" if overwrite else "x", + encoding="utf8", + ) as f: + f.write("# Generated by utils.py\n") + yaml.dump( + yaml_details, + f, + allow_unicode=True, + ) + except FileExistsError: + err.append(file_name) + + if len(err) > 0: + raise FileExistsError( + "Files were not created because they already exist (use --overwrite flag):" + f" {', '.join(err)}" + ) + + +def main() -> None: + """Parse CLI args and generate language-specific yaml files.""" + parser = argparse.ArgumentParser() + parser.add_argument( + "--overwrite", + default=True, + action="store_true", + help="Overwrite files if they already exist", + ) + parser.add_argument( + "--output-dir", + default="./direct", + help="Directory to write yaml files to", + ) + parser.add_argument( + "--mode", + default="prompt_4", + choices=["prompt_1", "prompt_2", "prompt_3", "prompt_4", "prompt_5"], + help="Prompt number", + ) + args = parser.parse_args() + + gen_lang_yamls(output_dir=args.output_dir, overwrite=args.overwrite, mode=args.mode) + + +if __name__ == "__main__": + main() diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_amh.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_amh.yaml new file mode 100644 index 0000000000000000000000000000000000000000..aaaaa6b8b653a2093459dfc1d7649c932b1e1957 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_amh.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: amh +include: afrimmlu_translate +task: afrimmlu_translate_amh_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..45298a1f4ed4c0f101f816b7adc46115da2b3aba --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrimmlu_translate +task: afrimmlu_translate_ewe_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ibo.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ibo.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6fe139910f24c12612a64c9633fd2dd625c580cc --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_ibo.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ibo +include: afrimmlu_translate +task: afrimmlu_translate_ibo_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c689952307fc07588af39a3c9db587aac74c4389 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrimmlu_translate +task: afrimmlu_translate_kin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..e245d7bdd3b96e16b84da7db505d0b005cb165fc --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_translate +task: afrimmlu_translate_lin_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..722ee9526102ab230cab4472842fda9e902f9119 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sna.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sna +include: afrimmlu_translate +task: afrimmlu_translate_sna_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..4e8893aa9d996552eb8a830dbaca3c8ad8988ee3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_sot.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: sot +include: afrimmlu_translate +task: afrimmlu_translate_sot_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..eb89697c67df29132bfb5db2c4d381f24f133e92 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_translate +task: afrimmlu_translate_swa_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_twi.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_twi.yaml new file mode 100644 index 0000000000000000000000000000000000000000..d672f6c768b823615c67ee46c9b48ae87d8177a9 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_twi.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: twi +include: afrimmlu_translate +task: afrimmlu_translate_twi_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_yor.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_yor.yaml new file mode 100644 index 0000000000000000000000000000000000000000..0f5eb7de392b44c98f4d415dff01eb0dda357247 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_yor.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: yor +include: afrimmlu_translate +task: afrimmlu_translate_yor_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..ae04b652fe227cecb10f25f8dd998a9ccb50dba3 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_1/afrimmlu_translate_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_translate +task: afrimmlu_translate_zul_prompt_1 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate new file mode 100644 index 0000000000000000000000000000000000000000..7a974279a3918de90369c391b09de818cb1b483d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate @@ -0,0 +1,32 @@ +tag: afrimmlu_tt_tasks +dataset_path: masakhane/afrimmlu-translate-test +dataset_name: null +output_type: multiple_choice +test_split: test +doc_to_text: !function utils.doc_to_text +doc_to_target: "{{['A', 'B', 'C', 'D'].index(answer)}}" +doc_to_choice: !function utils.doc_to_choice +should_decontaminate: true +doc_to_decontamination_query: "Question: {{question}}\nAnswer:" +metric_list: + - metric: f1 + aggregation: !function utils.weighted_f1_score + # aggregation: mean + average: weighted + hf_evaluate: true + higher_is_better: True + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" + - metric: acc + aggregation: mean + higher_is_better: true + ignore_case: true + ignore_punctuation: true + regexes_to_ignore: + - "," + - "\\$" +metadata: + version: 1.0 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_lin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_lin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..54163d5ccfad9c3973265750ca52b522cf685c7d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_lin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: lin +include: afrimmlu_translate +task: afrimmlu_translate_lin_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f267b6d087b15b710a44e8015f2d77070c35853d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_swa.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: swa +include: afrimmlu_translate +task: afrimmlu_translate_swa_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7f55271270b73c6c10eb29765ae8852383a2c832 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimmlu_translate +task: afrimmlu_translate_xho_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2ff80402b4d66972d6cbde5aec2eb63163aa35d5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_2/afrimmlu_translate_zul.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: zul +include: afrimmlu_translate +task: afrimmlu_translate_zul_prompt_2 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6f95b375edb041fd68f907a0b00411beb4808e50 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrimmlu_translate +task: afrimmlu_translate_orm_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_xho.yaml b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_xho.yaml new file mode 100644 index 0000000000000000000000000000000000000000..18678da825da9e065ed6825d3d7955e30e9c7fd0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrimmlu/translate/prompt_3/afrimmlu_translate_xho.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: xho +include: afrimmlu_translate +task: afrimmlu_translate_xho_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_2/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_2/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..5d1ac19e19b2e855c957e75f1c778366dfbc7e55 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_2/utils.py @@ -0,0 +1,6 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_target(doc): + replacements = {0: "True", 1: "Neither", 2: "False"} + return replacements[doc["label"]] diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1bd0829bfaf89354c5814eeffe1d4de8432fa540 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_kin.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: kin +doc_to_text: "Given the following premise and hypothesis in Kinyarwanda, identify\ + \ if the premise entails, contradicts, or is neutral towards the hypothesis. Please\ + \ respond with exact 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_kin_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..37a6d843e51870f6ad2845e1f385b6aacef2fab2 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_orm.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: orm +doc_to_text: "Given the following premise and hypothesis in Oromo, identify if the\ + \ premise entails, contradicts, or is neutral towards the hypothesis. Please respond\ + \ with exact 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_orm_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sna.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sna.yaml new file mode 100644 index 0000000000000000000000000000000000000000..c7e0f0b05000c40fbebca19921a2486a57d8054a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sna.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: sna +doc_to_text: "Given the following premise and hypothesis in chiShona, identify if\ + \ the premise entails, contradicts, or is neutral towards the hypothesis. Please\ + \ respond with exact 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_sna_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sot.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sot.yaml new file mode 100644 index 0000000000000000000000000000000000000000..0c0ccd9e64ee0b807fef742c873126666327f276 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_sot.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: sot +doc_to_text: "Given the following premise and hypothesis in Sesotho, identify if the\ + \ premise entails, contradicts, or is neutral towards the hypothesis. Please respond\ + \ with exact 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_sot_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..83b87141b4021baad581be7d6b60375c40ef73c5 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/afrixnli_zul.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: zul +doc_to_text: "Given the following premise and hypothesis in Zulu, identify if the\ + \ premise entails, contradicts, or is neutral towards the hypothesis. Please respond\ + \ with exact 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_zul_prompt_3 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/utils.py b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/utils.py new file mode 100644 index 0000000000000000000000000000000000000000..422ed169bffa2777cd91307c3ab097619e4d5399 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_3/utils.py @@ -0,0 +1,6 @@ +from lm_eval.utils import weighted_f1_score + + +def doc_to_target(doc): + replacements = {0: "entailment", 1: "neutral", 2: "contradiction"} + return replacements[doc["label"]] diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_eng.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_eng.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1ecb06d10497274dde56ab73302525add553254a --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_eng.yaml @@ -0,0 +1,9 @@ +# Generated by utils.py +dataset_name: eng +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the English language.\nAnalyze the premise and hypothesis given in English, and\ + \ determine the relationship between them.\n Respond with one of the following options:\ + \ 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis:\ + \ {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_eng_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..64157b549956f932b3ad5bd2f67610378d683596 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_ewe.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: ewe +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Ewe language.\nAnalyze the premise and hypothesis given in Ewe, and determine\ + \ the relationship between them.\n Respond with one of the following options: 'entailment',\ + \ 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_ewe_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..78da10cf7e04482eef5c0f9517a176196681c103 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_fra.yaml @@ -0,0 +1,9 @@ +# Generated by utils.py +dataset_name: fra +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the French language.\nAnalyze the premise and hypothesis given in French, and\ + \ determine the relationship between them.\n Respond with one of the following options:\ + \ 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis:\ + \ {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_fra_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f975d82b4a3b2ee0898aa8cc9aec225b3bb26e2d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_kin.yaml @@ -0,0 +1,9 @@ +# Generated by utils.py +dataset_name: kin +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Kinyarwanda language.\nAnalyze the premise and hypothesis given in Kinyarwanda,\ + \ and determine the relationship between them.\n Respond with one of the following\ + \ options: 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}}\ + \ \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_kin_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_lug.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_lug.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1553c620009ec7374378e584e5f7523ff6d57306 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_lug.yaml @@ -0,0 +1,9 @@ +# Generated by utils.py +dataset_name: lug +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Luganda language.\nAnalyze the premise and hypothesis given in Luganda, and\ + \ determine the relationship between them.\n Respond with one of the following options:\ + \ 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis:\ + \ {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_lug_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_swa.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_swa.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1c28aaae79accde1d796fb5dff56f8998130df0b --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_swa.yaml @@ -0,0 +1,9 @@ +# Generated by utils.py +dataset_name: swa +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Swahili language.\nAnalyze the premise and hypothesis given in Swahili, and\ + \ determine the relationship between them.\n Respond with one of the following options:\ + \ 'entailment', 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis:\ + \ {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_swa_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..6b535bc2d45555821810abb91755fb2afbae9bd1 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_wol.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: wol +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Wolof language.\nAnalyze the premise and hypothesis given in Wolof, and determine\ + \ the relationship between them.\n Respond with one of the following options: 'entailment',\ + \ 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_wol_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..1b4a232e395e36f6180a81247b23b624efcdbd05 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_4/afrixnli_zul.yaml @@ -0,0 +1,8 @@ +# Generated by utils.py +dataset_name: zul +doc_to_text: "You are an expert in Natural Language Inference (NLI) specializing in\ + \ the Zulu language.\nAnalyze the premise and hypothesis given in Zulu, and determine\ + \ the relationship between them.\n Respond with one of the following options: 'entailment',\ + \ 'contradiction', or 'neutral'. \n\nPremise: {{premise}} \nHypothesis: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_zul_prompt_4 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..7f60db0bffdd4ed7c367a10c6e365707a02348a4 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_ewe.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: ewe +doc_to_text: "Based on the given statement, is the following claim 'true', 'false',\ + \ or 'inconclusive'. \nStatement: {{premise}} \nClaim: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_ewe_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..13a8845cf1ca22f4a25c79931d526a3305a2172c --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_kin.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: kin +doc_to_text: "Based on the given statement, is the following claim 'true', 'false',\ + \ or 'inconclusive'. \nStatement: {{premise}} \nClaim: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_kin_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..f7f555db795996ac482deaae924db8af58e5c123 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_orm.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: orm +doc_to_text: "Based on the given statement, is the following claim 'true', 'false',\ + \ or 'inconclusive'. \nStatement: {{premise}} \nClaim: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_orm_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_zul.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_zul.yaml new file mode 100644 index 0000000000000000000000000000000000000000..2aa872b0410f56ea9e7ea19c4fb3d5adf93d323d --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/direct/prompt_5/afrixnli_zul.yaml @@ -0,0 +1,6 @@ +# Generated by utils.py +dataset_name: zul +doc_to_text: "Based on the given statement, is the following claim 'true', 'false',\ + \ or 'inconclusive'. \nStatement: {{premise}} \nClaim: {{hypothesis}}" +include: afrixnli_yaml +task: afrixnli_zul_prompt_5 diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_ewe.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_ewe.yaml new file mode 100644 index 0000000000000000000000000000000000000000..fe2fce97e33d8958a7a064aa25baa7e86d6f8f21 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_ewe.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: ewe +include: afrixnli_manual_direct_yaml +task: afrixnli_manual_direct_ewe diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_fra.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_fra.yaml new file mode 100644 index 0000000000000000000000000000000000000000..07c2f66238939ab19cce5f697826a5f53cdbe876 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_fra.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: fra +include: afrixnli_manual_direct_yaml +task: afrixnli_manual_direct_fra diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_kin.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_kin.yaml new file mode 100644 index 0000000000000000000000000000000000000000..611f61df85e89324769b6065e269e48ff3902190 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_kin.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: kin +include: afrixnli_manual_direct_yaml +task: afrixnli_manual_direct_kin diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_orm.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_orm.yaml new file mode 100644 index 0000000000000000000000000000000000000000..4931dc0a9ef3a58d9e9cdca3c6ab128333f7d3a0 --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_orm.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: orm +include: afrixnli_manual_direct_yaml +task: afrixnli_manual_direct_orm diff --git a/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_wol.yaml b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_wol.yaml new file mode 100644 index 0000000000000000000000000000000000000000..9f189d3975a07b125581d482e268205793e1577e --- /dev/null +++ b/lm-evaluation-harness/lm_eval/tasks/afrixnli/lai prompt/direct/afrixnli_manual_direct_wol.yaml @@ -0,0 +1,4 @@ +# Generated by utils.py +dataset_name: wol +include: afrixnli_manual_direct_yaml +task: afrixnli_manual_direct_wol