Pidoxy/cloud-notebook-runtime-blind-spots / evaluation_prompts_all.csv
Pidoxy's picture
download
raw
14.5 kB
case_id,platform,event,prompt
case_09,kaggle_interactive,midrun_feedback,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: I need to inspect validation loss and sample predictions while a long experiment is still running.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_17,kaggle_background,save_vs_checkpoint,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_01,local_jupyter,close_and_sleep,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_19,colab_interactive,close_and_sleep,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_12,kaggle_interactive,two_stage_workflow,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_10,kaggle_interactive,resume_after_crash,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_14,kaggle_background,internet_loss,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: My internet connection may disappear for one hour while the model is training.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_02,local_jupyter,internet_loss,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: My internet connection may disappear for one hour while the model is training.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_22,colab_interactive,resume_after_crash,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_06,local_jupyter,two_stage_workflow,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_03,local_jupyter,midrun_feedback,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: I need to inspect validation loss and sample predictions while a long experiment is still running.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_13,kaggle_background,close_and_sleep,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_16,kaggle_background,resume_after_crash,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_04,local_jupyter,resume_after_crash,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_05,local_jupyter,save_vs_checkpoint,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook.
Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_23,colab_interactive,save_vs_checkpoint,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_18,kaggle_background,two_stage_workflow,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_21,colab_interactive,midrun_feedback,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: I need to inspect validation loss and sample predictions while a long experiment is still running.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_24,colab_interactive,two_stage_workflow,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_08,kaggle_interactive,internet_loss,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: My internet connection may disappear for one hour while the model is training.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_11,kaggle_interactive,save_vs_checkpoint,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_15,kaggle_background,midrun_feedback,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.
Situation: I need to inspect validation loss and sample predictions while a long experiment is still running.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_20,colab_interactive,internet_loss,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime.
Situation: My internet connection may disappear for one hour while the model is training.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."
case_07,kaggle_interactive,close_and_sleep,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session.
Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep.
Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy."

Xet Storage Details

Size:
14.5 kB
·
Xet hash:
aea313b43eaa7bdf81e5947b10388e20f8f604046ee38d284ca0ed36c500124a

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.