Buckets:
| case_id,platform,event,prompt,model_id,seed,response,elapsed_seconds | |
| case_09,kaggle_interactive,midrun_feedback,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: I need to inspect validation loss and sample predictions while a long experiment is still running. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook runtime (client-side, browser-based)"", | |
| ""expected_behavior"": ""Validation loss and predictions can be inspected in real time via the notebook's output and visualization panels during execution, as long as the model is running in the same session."", | |
| ""main_risk"": ""Delayed or lost updates if the notebook session times out or the model runs for extended periods without explicit polling."", | |
| ""recommended_steps"": [ | |
| ""Monitor output logs and plots in real time during execution."", | |
| ""Use `print()` or `display()` statements to log validation metrics periodically."", | |
| ""Ensure the model runs in a session with persistent execution (e.g., via a kernel that doesn't auto-restart)."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's session timeout policies may affect real-time monitoring; behavior may vary with future updates to runtime stability."" | |
| }",18.52 | |
| case_17,kaggle_background,save_vs_checkpoint,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Remote (Kaggle backend)"", | |
| ""expected_behavior"": ""Saving a notebook via 'Save & Run All' does not automatically preserve training state or model progress; it only saves the notebook's code and environment configuration."", | |
| ""main_risk"": ""Loss of model progress if training was not explicitly stopped or saved via model checkpointing or persistent storage."", | |
| ""recommended_steps"": [ | |
| ""Explicitly save model checkpoints using `model.save()` or similar methods."", | |
| ""Use persistent storage (e.g., Google Drive, local file system) for model artifacts."", | |
| ""Verify training state is saved independently of notebook saves."" | |
| ], | |
| ""policy_uncertainty"": ""No, the platform does not guarantee automatic persistence of training state beyond the notebook save."" | |
| }",15.22 | |
| case_01,local_jupyter,close_and_sleep,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""local laptop"", | |
| ""expected_behavior"": ""The Python kernel continues running on the laptop; Jupyter will not terminate the process when the browser tab is closed."", | |
| ""main_risk"": ""The kernel may be interrupted by system sleep or power-off if not explicitly managed."", | |
| ""recommended_steps"": [ | |
| ""Keep the laptop powered and awake during training."", | |
| ""Use `systemd` or `screen` to run the kernel in the background if possible."", | |
| ""Enable 'sleep' or 'hibernate' settings to avoid interruption."" | |
| ], | |
| ""policy_uncertainty"": ""Jupyter's behavior when the browser is closed is not guaranteed by platform policy and may vary with updates."" | |
| }",14.2 | |
| case_19,colab_interactive,close_and_sleep,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime"", | |
| ""expected_behavior"": ""The model training continues on the Colab instance; the browser tab closure does not terminate the process, as long as the runtime is not explicitly stopped or killed."", | |
| ""main_risk"": ""The runtime may be terminated due to inactivity or resource limits, especially if no activity occurs for more than 24 hours."", | |
| ""recommended_steps"": [ | |
| ""Keep the Colab runtime active by periodically checking the notebook or using 'Runtime > Keep Running' if available."", | |
| ""Use a persistent Google Cloud VM or a local machine for long-running jobs if stability is critical."", | |
| ""Set up a background job with a cloud-based ML platform (e.g., Vertex AI) with explicit job scheduling."" | |
| ], | |
| ""policy_uncertainty"": ""Google Colab's auto-termination policy for idle runtimes may change without notice."" | |
| }",17.7 | |
| case_12,kaggle_interactive,two_stage_workflow,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook session (client-side browser, local runtime)"", | |
| ""expected_behavior"": ""Initial runs execute quickly in browser; errors surface immediately. Full experiments require explicit deployment to a remote server or cloud instance."", | |
| ""main_risk"": ""Browser-based execution may fail silently or misreport due to resource limits or connectivity issues."", | |
| ""recommended_steps"": [ | |
| ""Run small, isolated tests first to validate logic and catch syntax errors."", | |
| ""Use Kaggle's 'Submit' or 'Deploy' feature to offload to a reliable remote runtime for unattended execution."", | |
| ""Verify logs and output in both environments to ensure consistency."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's runtime policies may change regarding long-running or resource-intensive experiments."" | |
| }",15.57 | |
| case_10,kaggle_interactive,resume_after_crash,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook session (cloud-based, remote runtime)"", | |
| ""expected_behavior"": ""Training can be paused and resumed via checkpointing; Kaggle supports resuming from saved checkpoints without restarting from epoch 1."", | |
| ""main_risk"": ""Loss of progress if the runtime is terminated or the session expires before resuming."", | |
| ""recommended_steps"": [ | |
| ""Save model checkpoints every few epochs using `torch.save()` or `tf.keras.Model.save()`."", | |
| ""Pause training, then resume from the last saved checkpoint using the same model and data."", | |
| ""Ensure the model state and optimizer are properly serialized and restored."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's support for resuming training may change with platform updates; verify via Kaggle's official documentation."" | |
| }",15.82 | |
| case_14,kaggle_background,internet_loss,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: My internet connection may disappear for one hour while the model is training. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Remote (Kaggle backend)"", | |
| ""expected_behavior"": ""Training continues uninterrupted if the connection is lost; Kaggle resumes execution upon reconnection."", | |
| ""main_risk"": ""Training job may be interrupted or fail if the remote runtime is disconnected for >15 minutes without recovery."", | |
| ""recommended_steps"": [ | |
| ""Save the notebook with 'Save & Run All' before disconnecting."", | |
| ""Set up a checkpointing strategy to save model progress periodically."", | |
| ""Use Kaggle's 'Resume' feature to restart training after reconnection."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's automatic recovery behavior during network outages is not publicly documented and may change."" | |
| }",14.08 | |
| case_02,local_jupyter,internet_loss,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: My internet connection may disappear for one hour while the model is training. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""local laptop"", | |
| ""expected_behavior"": ""Training continues locally; no internet required once computation starts."", | |
| ""main_risk"": ""Training interruption due to lost internet, potentially leading to incomplete or corrupted model."", | |
| ""recommended_steps"": [ | |
| ""Start training with a local dataset to avoid internet dependency."", | |
| ""Save checkpoints every 5-10 minutes to allow recovery on restart."", | |
| ""Ensure the model uses persistent storage (e.g., disk) for intermediate results."", | |
| ""Set up a local backup or auto-recovery mechanism for training state."" | |
| ], | |
| ""policy_uncertainty"": ""None. Behavior is fully determined by local Python execution and file system."" | |
| }",13.65 | |
| case_22,colab_interactive,resume_after_crash,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime"", | |
| ""expected_behavior"": ""Training can be paused and resumed via checkpointing; Colab supports resuming from saved checkpoints without restarting from epoch 1."", | |
| ""main_risk"": ""Data or model state corruption if checkpoints are not properly saved or restored."", | |
| ""recommended_steps"": [ | |
| ""Save model state and training progress to a local file at regular intervals."", | |
| ""Use `torch.save()` or `tf.keras.models.save_model()` to store checkpoints."", | |
| ""Resume training from the last checkpoint using the same model and optimizer state."" | |
| ], | |
| ""policy_uncertainty"": ""Colab's long-running session support may change; verify via official documentation."" | |
| }",13.85 | |
| case_06,local_jupyter,two_stage_workflow,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""local laptop via Jupyter Notebook"", | |
| ""expected_behavior"": ""Quick error detection during development; full experiments run reliably in unattended mode via script or CI/CD."", | |
| ""main_risk"": ""Local environment drift or uncaught bugs during unattended execution due to missing monitoring."", | |
| ""recommended_steps"": [ | |
| ""Run experiments in Jupyter with immediate error logging and debugging."", | |
| ""Automate full experiments using Python scripts with logging and error trapping."", | |
| ""Use version control to track experiment configurations and code changes."", | |
| ""Set up a CI/CD pipeline for unattended execution with notifications on failure."" | |
| ], | |
| ""policy_uncertainty"": ""None. Behavior is technically verifiable and independent of platform policy."" | |
| }",14.39 | |
| case_03,local_jupyter,midrun_feedback,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: I need to inspect validation loss and sample predictions while a long experiment is still running. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Local laptop via Jupyter Notebook"", | |
| ""expected_behavior"": ""Validation loss and predictions can be inspected in real time using Jupyter's live output and model evaluation functions, as long as the training loop is accessible and data is loaded."", | |
| ""main_risk"": ""Data leakage or overfitting due to unmonitored validation during long runs without proper data splitting or early stopping."", | |
| ""recommended_steps"": [ | |
| ""Add real-time logging of validation loss to a file or notebook output."", | |
| ""Use a separate validation set and ensure it's not used during training."", | |
| ""Implement periodic model checkpointing and prediction sampling."" | |
| ], | |
| ""policy_uncertainty"": ""None. Behavior is fully determined by standard ML practices and Jupyter's execution model."" | |
| }",15.22 | |
| case_13,kaggle_background,close_and_sleep,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""remote (Kaggle backend)"", | |
| ""expected_behavior"": ""The notebook continues training on Kaggle's infrastructure; the browser tab closure does not interrupt execution."", | |
| ""main_risk"": ""Unintended interruption due to platform policy changes or unexpected timeouts (not guaranteed)."", | |
| ""recommended_steps"": [ | |
| ""Save and run the notebook with 'Save & Run All' to initiate remote execution."", | |
| ""Close the browser tab; training proceeds independently on Kaggle's servers."", | |
| ""Monitor progress via the notebook's status dashboard, not the browser."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle may alter job timeout policies or remote execution behavior without notice."" | |
| }",13.84 | |
| case_16,kaggle_background,resume_after_crash,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle remote runtime (cloud-based)"", | |
| ""expected_behavior"": ""Training resumes from the last saved checkpoint if the runtime is interrupted and the model state is preserved via checkpointing."", | |
| ""main_risk"": ""Loss of progress if the remote runtime is terminated or crashes without saving state."", | |
| ""recommended_steps"": [ | |
| ""Enable automatic checkpointing during training (e.g., save every 100 epochs)."", | |
| ""Use Kaggle's 'Save & Run All' with explicit checkpointing in code (e.g., `torch.save(model.state_dict(), 'checkpoint.pth')`)."", | |
| ""Verify that the model state is saved in a persistent storage location (e.g., local disk or cloud bucket)."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's support for resuming training from checkpoints after runtime interruption is not officially documented and may change."" | |
| }",17.47 | |
| case_04,local_jupyter,resume_after_crash,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: The training process could fail after six hours, and I need to continue without restarting from epoch one. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""local laptop via Jupyter Notebook"", | |
| ""expected_behavior"": ""Training can be paused and resumed using model checkpoints and state persistence; no restart from epoch 1 is required if the model state is saved."", | |
| ""main_risk"": ""Data corruption or model state loss if checkpoints are not properly saved or restored."", | |
| ""recommended_steps"": [ | |
| ""Save model state and optimizer state at regular intervals using `torch.save()` or `model.save_pretrained()`."", | |
| ""Use a checkpointing strategy (e.g., every 1000 steps) to preserve training progress."", | |
| ""Resume training from the last checkpoint using the saved state upon restart."" | |
| ], | |
| ""policy_uncertainty"": ""None. The behavior is technically verifiable and independent of platform policy."" | |
| }",15.59 | |
| case_05,local_jupyter,save_vs_checkpoint,"I am running a machine-learning experiment. The Python kernel is running on my own laptop through Jupyter Notebook. | |
| Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""local laptop via Jupyter Notebook"", | |
| ""expected_behavior"": ""Saving a notebook does not automatically save model training state or progress; only the notebook's code and metadata are preserved."", | |
| ""main_risk"": ""Loss of training progress if model state is not explicitly saved to disk or a checkpoint file."", | |
| ""recommended_steps"": [ | |
| ""Explicitly save model and training state using `model.save()` or `torch.save()`/`joblib.dump()`"", | |
| ""Use version control (e.g., Git) to track model checkpoints and training logs"", | |
| ""Enable automatic checkpointing in training loops"" | |
| ], | |
| ""policy_uncertainty"": ""No"" | |
| }",13.4 | |
| case_23,colab_interactive,save_vs_checkpoint,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime"", | |
| ""expected_behavior"": ""Saving the notebook does not automatically save training state or model progress; only the notebook's code and environment configuration are preserved."", | |
| ""main_risk"": ""Loss of model progress if training state is not explicitly saved to disk or cloud storage."", | |
| ""recommended_steps"": [ | |
| ""Explicitly save model checkpoints using `torch.save()` or `tf.keras.models.save_model()`"", | |
| ""Store trained models in Google Drive or a cloud bucket via `gdown` or `gcs`"", | |
| ""Use version control (e.g., Git) to track code and model artifacts"" | |
| ], | |
| ""policy_uncertainty"": ""No, saving the notebook does not preserve training state; behavior depends on explicit model saving actions."" | |
| }",15.58 | |
| case_18,kaggle_background,two_stage_workflow,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Remote (Kaggle backend)"", | |
| ""expected_behavior"": ""Save & Run All triggers a remote job; initial runs provide quick feedback, full experiments run unattended."", | |
| ""main_risk"": ""Delayed error detection due to remote execution latency or uncaught failures in background jobs."", | |
| ""recommended_steps"": [ | |
| ""Run small, isolated tests locally first to validate logic."", | |
| ""Use Kaggle's 'Run' button for immediate feedback on small batches."", | |
| ""Enable logging and error notifications for remote jobs."", | |
| ""Set up automated email or webhook alerts for job failures."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's remote job reliability and error notification policies may change without notice."" | |
| }",14.4 | |
| case_21,colab_interactive,midrun_feedback,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: I need to inspect validation loss and sample predictions while a long experiment is still running. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime (client-side, in-browser)"", | |
| ""expected_behavior"": ""Validation loss and predictions can be inspected in real time via notebook cells or visualization tools during execution, as long as the model is running in the same runtime."", | |
| ""main_risk"": ""Data leakage or incorrect validation if the model is not properly split or evaluated in a separate validation loop."", | |
| ""recommended_steps"": [ | |
| ""Monitor validation loss in the notebook's output logs or plots."", | |
| ""Use `model.evaluate()` with validation data in a separate cell to inspect metrics."", | |
| ""Ensure validation data is not used during training to avoid data leakage."" | |
| ], | |
| ""policy_uncertainty"": ""No known platform policy changes affect real-time validation monitoring in Colab; behavior is consistent with standard ML workflows."" | |
| }",16.06 | |
| case_24,colab_interactive,two_stage_workflow,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: I need quick feedback to catch errors first, followed by a reliable unattended full experiment. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime (interactive, browser-based)"", | |
| ""expected_behavior"": ""Quick error detection during initial execution; full experiment runs reliably in unattended mode via persistent backend."", | |
| ""main_risk"": ""Transient runtime failures due to browser session timeouts or network instability."", | |
| ""recommended_steps"": [ | |
| ""Run code in interactive mode to catch syntax or logic errors immediately"", | |
| ""Save experiment configuration and use Colab's 'Runtime > Run all' for unattended execution"", | |
| ""Enable persistent storage via Google Drive or local file system for reproducibility"" | |
| ], | |
| ""policy_uncertainty"": ""Colab's long-term runtime availability and auto-restart policies may change without notice."" | |
| }",14.11 | |
| case_08,kaggle_interactive,internet_loss,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: My internet connection may disappear for one hour while the model is training. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook session (local runtime)"", | |
| ""expected_behavior"": ""Training continues locally until interrupted by a timeout or user action; no automatic recovery from disconnection."", | |
| ""main_risk"": ""Data loss or incomplete training due to unhandled disconnection without explicit checkpointing."", | |
| ""recommended_steps"": [ | |
| ""Enable automatic checkpointing using `torch.save()` or `joblib` to save model state periodically."", | |
| ""Use a persistent storage location (e.g., Google Drive) to back up training progress."", | |
| ""Set up a timeout or interrupt handler to pause training and save state if the session disconnects."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's automatic session recovery or remote reconnection policies are not publicly documented and may change."" | |
| }",15.56 | |
| case_11,kaggle_interactive,save_vs_checkpoint,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook session (cloud-based)"", | |
| ""expected_behavior"": ""Saving the notebook saves the code and environment state but does not automatically preserve training state or model progress unless explicitly saved via model checkpointing or persistent storage."", | |
| ""main_risk"": ""Loss of training progress if no explicit model checkpointing or persistent storage was used."", | |
| ""recommended_steps"": [ | |
| ""Enable model checkpointing during training using `model.save()` or similar methods."", | |
| ""Save trained models to a persistent location (e.g., Google Drive, local file system)."", | |
| ""Verify that training state is explicitly saved, not just the notebook code."" | |
| ], | |
| ""policy_uncertainty"": ""No, the platform does not guarantee automatic persistence of training state beyond the notebook session."" | |
| }",15.66 | |
| case_15,kaggle_background,midrun_feedback,"I am running a machine-learning experiment. I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job. | |
| Situation: I need to inspect validation loss and sample predictions while a long experiment is still running. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Remote (Kaggle backend)"", | |
| ""expected_behavior"": ""Validation loss and predictions are not visible in real-time during execution; only final results are accessible post-completion."", | |
| ""main_risk"": ""Misleading assumption that live monitoring is possible due to platform UI limitations."", | |
| ""recommended_steps"": [ | |
| ""Check the notebook's output logs for validation metrics during execution."", | |
| ""Use Kaggle's 'Run All' logs to verify if validation loss is logged."", | |
| ""Set up real-time logging via custom scripts or external monitoring tools."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's real-time validation visibility policy may change without notice."" | |
| }",13.54 | |
| case_20,colab_interactive,internet_loss,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Google Colab runtime. | |
| Situation: My internet connection may disappear for one hour while the model is training. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Google Colab runtime"", | |
| ""expected_behavior"": ""Training continues locally until interrupted by a timeout or manual stop; no automatic recovery or checkpointing is guaranteed."", | |
| ""main_risk"": ""Data loss or incomplete training due to unattended runtime termination without explicit save or checkpointing."", | |
| ""recommended_steps"": [ | |
| ""Save model checkpoints every 10 minutes using `torch.save()` or `tf.keras.models.save_model()`."", | |
| ""Enable persistent storage via Google Drive to preserve training state."", | |
| ""Set up a manual interrupt and resume mechanism using `try-except` and `RuntimeError` checks."", | |
| ""Run training in a separate, long-lived environment if possible."" | |
| ], | |
| ""policy_uncertainty"": ""Google Colab may terminate runtimes without notice; no official policy guarantees uptime or automatic recovery."" | |
| }",16.58 | |
| case_07,kaggle_interactive,close_and_sleep,"I am running a machine-learning experiment. I started the code by clicking Run in an interactive Kaggle notebook session. | |
| Situation: A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep. | |
| Return one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.",Qwen/Qwen3-4B-Instruct-2507,42,"{ | |
| ""computation_location"": ""Kaggle notebook (client-side browser)"", | |
| ""expected_behavior"": ""The model training runs locally on the user's laptop; Kaggle's backend does not automatically resume or manage long-running jobs after browser closure."", | |
| ""main_risk"": ""Data loss or incomplete training if the session is terminated without proper saving or checkpointing."", | |
| ""recommended_steps"": [ | |
| ""Save model checkpoints periodically during training."", | |
| ""Use Kaggle's 'Save' feature to store notebook state."", | |
| ""Ensure the training script includes a stop condition and logs progress."" | |
| ], | |
| ""policy_uncertainty"": ""Kaggle's support for background execution or auto-resume of long-running jobs is not officially documented or guaranteed."" | |
| }",14.48 | |
Xet Storage Details
- Size:
- 34.5 kB
- Xet hash:
- 76af0f3ee201e696b3367cc1d6eff71f51aa4f312db3974fce00fb3940b0ff66
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.