Buckets:
| {"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[],"dockerImageVersionId":28755,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# 1. Reproducible environment and model setup\n# This cell installs only what is needed, checks the GPU, and loads one eligible open-weight model.\n\nimport sys, subprocess, importlib.util\nfrom importlib.metadata import version, PackageNotFoundError\nfrom packaging.version import parse as vparse\n\nrequired = {\n \"transformers\": \"4.51.0\",\n \"accelerate\": \"0.30.0\",\n \"bitsandbytes\": \"0.43.0\",\n}\n\nto_install = []\nfor package, minimum in required.items():\n try:\n if vparse(version(package)) < vparse(minimum):\n to_install.append(f\"{package}>={minimum}\")\n except PackageNotFoundError:\n to_install.append(f\"{package}>={minimum}\")\n\nif to_install:\n print(\"Installing:\", to_install)\n subprocess.check_call([sys.executable, \"-m\", \"pip\", \"install\", \"-q\", \"-U\", *to_install])\n\nimport random, json, time\nfrom pathlib import Path\nimport pandas as pd\nimport torch\nfrom transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig\n\nSEED = 42\nrandom.seed(SEED)\ntorch.manual_seed(SEED)\n\nMODEL_ID = \"Qwen/Qwen3-4B-Instruct-2507\"\nOUTPUT_DIR = Path(\"/kaggle/working/cloud_runtime_eval\")\nOUTPUT_DIR.mkdir(parents=True, exist_ok=True)\n\nassert torch.cuda.is_available(), \"GPU is not available. In Session options, select GPU T4 x2.\"\nprint(\"GPU:\", torch.cuda.get_device_name(0))\nprint(\"Model:\", MODEL_ID)\n\nquantization_config = BitsAndBytesConfig(\n load_in_4bit=True,\n bnb_4bit_quant_type=\"nf4\",\n bnb_4bit_compute_dtype=torch.float16,\n bnb_4bit_use_double_quant=True,\n)\n\ntokenizer = AutoTokenizer.from_pretrained(MODEL_ID)\nmodel = AutoModelForCausalLM.from_pretrained(\n MODEL_ID,\n device_map=\"auto\",\n quantization_config=quantization_config,\n torch_dtype=torch.float16,\n)\nmodel.eval()\nprint(\"Model loaded successfully.\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-09-23T10:55:24.644733Z","iopub.execute_input":"2026-09-23T10:55:24.645118Z","iopub.status.idle":"2026-09-23T10:56:43.40901Z","shell.execute_reply.started":"2026-09-23T10:55:24.645096Z","shell.execute_reply":"2026-09-23T10:56:43.408024Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Evaluating cloud-notebook execution-lifecycle reasoning\n\n## Applicant's hypothesis — complete this in your own words before the full run\n\n**My lived experience:** [While running a long machine-learning experiment on Kaggle, I was unsure whether the computation would continue if I left the webpage or my laptop hibernated. I initially treated the browser, my laptop, and Kaggle’s remote runtime as though they shared the same lifecycle. I later used Kaggle’s “Save & Run All” option so the notebook could execute as a background version, although this meant waiting for the cells to run sequentially before seeing all the outputs.]\n\n**My hypothesis:** [I suspect that small instruction models do not consistently distinguish between local computation, interactive cloud sessions, and background cloud jobs. They may consequently give generic advice, such as keeping the laptop awake, without explaining when browser disconnection matters, when a remote runtime continues independently, and why saving notebook code is different from checkpointing an experiment.]\n\n## Question being tested\n\nCan a small frontier instruction model correctly distinguish:\n\n- computation running on a local laptop from computation running on a remote service;\n- closing a browser or losing internet from termination of the remote runtime;\n- an interactive Kaggle session from a Kaggle **Save & Run All** background version;\n- saving notebook code from checkpointing model weights, optimizer state, and progress;\n- immediate interactive feedback from reproducible background execution?\n\n## Model choice\n\nThis notebook evaluates **Qwen/Qwen3-4B-Instruct-2507**. It is an ungated 4B-parameter instruction model under the Apache-2.0 license, falls inside the required 0.6–6B range, and is small enough to run in 4-bit form on a Kaggle T4 GPU.\n\n## Procedure\n\nThe same six events are tested across four execution environments. Prompt order is shuffled with a fixed seed. Generation is deterministic. Start with the four-case pilot, inspect it yourself, then switch to the full 24-case run. Responses are written to disk after every example so partial progress is not lost.","metadata":{}},{"cell_type":"code","source":"# 2. Build a controlled 6-event × 4-environment prompt matrix\n# You may edit the wording, but preserve the matched structure so comparisons remain fair.\n\nRUN_MODE = \"full\" # change to \"full\" only after you inspect the four pilot responses\nN_PILOT = 4\n\nplatforms = {\n \"local_jupyter\": \"The Python kernel is running on my own laptop through Jupyter Notebook.\",\n \"kaggle_interactive\": \"I started the code by clicking Run in an interactive Kaggle notebook session.\",\n \"kaggle_background\": \"I created a Kaggle notebook version using Save & Run All, which runs remotely as a background job.\",\n \"colab_interactive\": \"I started the code by clicking Run in an interactive Google Colab runtime.\",\n}\n\nevents = {\n \"close_and_sleep\": \"A model-training cell may take eight hours. I want to close the browser tab and let my laptop sleep.\",\n \"internet_loss\": \"My internet connection may disappear for one hour while the model is training.\",\n \"midrun_feedback\": \"I need to inspect validation loss and sample predictions while a long experiment is still running.\",\n \"resume_after_crash\": \"The training process could fail after six hours, and I need to continue without restarting from epoch one.\",\n \"save_vs_checkpoint\": \"I saved the notebook after editing its code. I need to know whether that also saved the current training state and model progress.\",\n \"two_stage_workflow\": \"I need quick feedback to catch errors first, followed by a reliable unattended full experiment.\",\n}\n\nrecords = []\ncase_id = 0\nfor platform_id, platform_text in platforms.items():\n for event_id, event_text in events.items():\n case_id += 1\n prompt = f\"\"\"I am running a machine-learning experiment. {platform_text}\n\nSituation: {event_text}\n\nReturn one valid JSON object with exactly these keys: computation_location, expected_behavior, main_risk, recommended_steps, and policy_uncertainty. recommended_steps must be an array. Use no more than 180 words in total. Distinguish verified technical behavior from anything that depends on current platform policy.\"\"\"\n records.append({\n \"case_id\": f\"case_{case_id:02d}\",\n \"platform\": platform_id,\n \"event\": event_id,\n \"prompt\": prompt,\n })\n\nprompt_df = pd.DataFrame(records).sample(frac=1, random_state=SEED).reset_index(drop=True)\nrun_df = prompt_df.head(N_PILOT).copy() if RUN_MODE == \"pilot\" else prompt_df.copy()\n\nprint(f\"Mode: {RUN_MODE} | cases selected: {len(run_df)} of {len(prompt_df)}\")\ndisplay(run_df[[\"case_id\", \"platform\", \"event\", \"prompt\"]])\n\nprompt_df.to_csv(OUTPUT_DIR / \"evaluation_prompts_all.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T10:56:43.410601Z","iopub.execute_input":"2026-09-23T10:56:43.411229Z","iopub.status.idle":"2026-09-23T10:56:43.517603Z","shell.execute_reply.started":"2026-09-23T10:56:43.411198Z","shell.execute_reply":"2026-09-23T10:56:43.517031Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Generate one deterministic response per selected prompt\n# A CSV and JSONL file are updated after every case.\n\nSYSTEM_MESSAGE = (\n \"You are a technically precise machine-learning systems assistant. \"\n \"Do not assume that a browser, laptop, and remote runtime share the same lifecycle. \"\n \"If behavior depends on current platform policy, say so rather than inventing a guarantee. Follow the requested JSON format without a preamble.\"\n)\n\ndef generate_response(prompt, max_new_tokens=300):\n messages = [\n {\"role\": \"system\", \"content\": SYSTEM_MESSAGE},\n {\"role\": \"user\", \"content\": prompt},\n ]\n text = tokenizer.apply_chat_template(\n messages,\n tokenize=False,\n add_generation_prompt=True,\n )\n inputs = tokenizer(text, return_tensors=\"pt\").to(model.device)\n with torch.inference_mode():\n output = model.generate(\n **inputs,\n max_new_tokens=max_new_tokens,\n do_sample=False,\n pad_token_id=tokenizer.eos_token_id,\n )\n generated = output[0, inputs[\"input_ids\"].shape[1]:]\n return tokenizer.decode(generated, skip_special_tokens=True).strip()\n\nresults = []\ncsv_path = OUTPUT_DIR / f\"raw_responses_{RUN_MODE}.csv\"\njsonl_path = OUTPUT_DIR / f\"raw_responses_{RUN_MODE}.jsonl\"\n\nfor position, row in run_df.iterrows():\n started = time.time()\n print(f\"[{position + 1}/{len(run_df)}] {row['case_id']} | {row['platform']} | {row['event']}\")\n response = generate_response(row[\"prompt\"])\n result = {\n **row.to_dict(),\n \"model_id\": MODEL_ID,\n \"seed\": SEED,\n \"response\": response,\n \"elapsed_seconds\": round(time.time() - started, 2),\n }\n results.append(result)\n\n # Save partial progress after every case.\n pd.DataFrame(results).to_csv(csv_path, index=False)\n with jsonl_path.open(\"w\", encoding=\"utf-8\") as f:\n for item in results:\n f.write(json.dumps(item, ensure_ascii=False) + \"\\n\")\n\n print(response)\n print(\"-\" * 100)\n\nresults_df = pd.DataFrame(results)\nprint(f\"Saved {len(results_df)} responses to {csv_path}\")\ndisplay(results_df[[\"case_id\", \"platform\", \"event\", \"response\"]])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T10:56:43.518587Z","iopub.execute_input":"2026-09-23T10:56:43.518947Z","iopub.status.idle":"2026-09-23T11:02:48.379425Z","shell.execute_reply.started":"2026-09-23T10:56:43.518908Z","shell.execute_reply":"2026-09-23T11:02:48.378839Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Human review — scores and observations supplied by the applicant\n# Use 1 = present/correct, 0 = absent/incorrect, and NA = not applicable.\nlabel_columns = [\n \"location_correct\",\n \"disconnect_vs_termination_correct\",\n \"interactive_vs_background_correct\",\n \"save_vs_checkpoint_correct\",\n \"platform_specific_actionable_advice\",\n \"states_policy_uncertainty\",\n \"overall_quality_0_to_3\",\n \"failure_description_in_my_own_words\",\n]\n\nreview_df = results_df.copy()\nfor column in label_columns:\n review_df[column] = \"\"\n\napplicant_reviews = {\n \"case_09\": {\n \"overall_quality_0_to_3\": 1,\n \"failure_description_in_my_own_words\": \"The response incorrectly describes Kaggle computation as client-side and browser-based.\",\n },\n \"case_17\": {\n \"overall_quality_0_to_3\": 2,\n \"failure_description_in_my_own_words\": \"The response correctly separates notebook saving from checkpointing, but suggests storage options that are not well tailored to Kaggle.\",\n },\n \"case_01\": {\n \"overall_quality_0_to_3\": 1,\n \"failure_description_in_my_own_words\": \"The response correctly identifies the local laptop as the computation location, but recommending sleep or hibernation contradicts the goal of preventing interruption.\",\n },\n \"case_19\": {\n \"overall_quality_0_to_3\": 2,\n \"failure_description_in_my_own_words\": \"The response is broadly correct, but the precise Colab time limit and the Keep Running option may be uncertain or outdated.\",\n },\n}\n\nfor case_id, values in applicant_reviews.items():\n for column, value in values.items():\n review_df.loc[review_df[\"case_id\"] == case_id, column] = value\n\nreview_path = OUTPUT_DIR / f\"human_review_{RUN_MODE}_FILL_THIS_YOURSELF.csv\"\nreview_df.to_csv(review_path, index=False)\nprint(\"Human-review sheet updated:\", review_path)\ndisplay(review_df[[\"case_id\", \"platform\", \"event\", \"overall_quality_0_to_3\", \"failure_description_in_my_own_words\"]])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T11:02:48.38091Z","iopub.execute_input":"2026-09-23T11:02:48.381257Z","iopub.status.idle":"2026-09-23T11:02:48.406847Z","shell.execute_reply.started":"2026-09-23T11:02:48.381226Z","shell.execute_reply":"2026-09-23T11:02:48.406193Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 5. Full-text review: every prompt and model response, without table truncation\nfrom pathlib import Path\nfrom html import escape\nfrom IPython.display import HTML, display\nimport pandas as pd\n\nreview_path = Path('/kaggle/working/cloud_runtime_eval/human_review_full_FILL_THIS_YOURSELF.csv')\nprint('Saved review file exists:', review_path.exists())\nif review_path.exists():\n review_df = pd.read_csv(review_path, keep_default_na=False)\n print(f'Full responses: {len(review_df)} cases. Scroll down to read each one.')\n for _, row in review_df.iterrows():\n score = row['overall_quality_0_to_3']\n note = row['failure_description_in_my_own_words']\n heading = f\"{row['case_id']} — {row['platform']} / {row['event']}\"\n display(HTML(\n \"<section style='border:1px solid #bbb;border-radius:8px;padding:16px;margin:20px 0;max-width:1000px'>\"\n f\"<h3>{escape(heading)}</h3>\"\n \"<strong>Full prompt</strong>\"\n f\"<pre style='white-space:pre-wrap;overflow-wrap:anywhere;font:inherit'>{escape(str(row['prompt']))}</pre>\"\n \"<strong>Full model response</strong>\"\n f\"<pre style='white-space:pre-wrap;overflow-wrap:anywhere;font:inherit'>{escape(str(row['response']))}</pre>\"\n f\"<p><strong>Your score:</strong> {escape(str(score)) if str(score) else 'Not yet reviewed'}<br>\"\n f\"<strong>Your note:</strong> {escape(str(note)) if str(note) else 'Not yet reviewed'}</p>\"\n \"</section>\"\n ))\nelse:\n print('The temporary Kaggle working directory was cleared when the session restarted. The previous run must be restored from a saved version or run again.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T11:02:48.4079Z","iopub.execute_input":"2026-09-23T11:02:48.408269Z","iopub.status.idle":"2026-09-23T11:02:48.453082Z","shell.execute_reply.started":"2026-09-23T11:02:48.408236Z","shell.execute_reply":"2026-09-23T11:02:48.45258Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 6. Create a backup archive before the temporary Kaggle session ends\nfrom pathlib import Path\nfrom zipfile import ZipFile, ZIP_DEFLATED\n\noutput_dir = Path('/kaggle/working/cloud_runtime_eval')\nfiles_to_keep = [\n output_dir / 'evaluation_prompts_all.csv',\n output_dir / 'raw_responses_full.csv',\n output_dir / 'raw_responses_full.jsonl',\n output_dir / 'human_review_full_FILL_THIS_YOURSELF.csv',\n]\nmissing = [str(p) for p in files_to_keep if not p.exists()]\nif missing:\n print('Missing files:', missing)\nelse:\n archive = output_dir / 'cloud_runtime_evaluation_backup.zip'\n with ZipFile(archive, 'w', ZIP_DEFLATED) as z:\n for path in files_to_keep:\n z.write(path, arcname=path.name)\n print('Backup ready:', archive)\n print('To download: in the right Output sidebar, expand /kaggle/working > cloud_runtime_eval, then use the ZIP file’s More actions > Download.')\n print('Create a fresh backup again after adding your own reviews.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T11:22:13.66076Z","iopub.execute_input":"2026-09-23T11:22:13.661665Z","iopub.status.idle":"2026-09-23T11:22:13.67163Z","shell.execute_reply.started":"2026-09-23T11:22:13.661632Z","shell.execute_reply":"2026-09-23T11:22:13.671039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 7. Applicant's review form — judgments must be yours, not generated by the notebook\nimport pandas as pd\nimport ipywidgets as W\nfrom html import escape\nfrom IPython.display import HTML, display, clear_output\nfrom pathlib import Path\n\nreview_path = Path('/kaggle/working/cloud_runtime_eval/human_review_full_FILL_THIS_YOURSELF.csv')\nassert review_path.exists(), 'Review CSV is missing. Restore it from your backup.'\nreview_data = pd.read_csv(review_path, keep_default_na=False)\ncase_ids = review_data['case_id'].tolist()\n\ncase_choice = W.Dropdown(options=case_ids, description='Case:', layout=W.Layout(width='400px'))\nscore_choice = W.Dropdown(options=[('Choose a score',''),('0 — incorrect','0'),('1 — mostly incorrect','1'),('2 — partly correct','2'),('3 — correct','3')], description='Score:', layout=W.Layout(width='400px'))\nnote_input = W.Textarea(placeholder='In your own words: what did the model get right or wrong?', description='Your note:', layout=W.Layout(width='900px', height='110px'))\ncriteria = {\n 'location_correct': 'Computation location',\n 'disconnect_vs_termination_correct': 'Disconnect vs termination',\n 'interactive_vs_background_correct': 'Interactive vs background',\n 'save_vs_checkpoint_correct': 'Save vs checkpoint',\n 'platform_specific_actionable_advice': 'Platform-specific advice',\n 'states_policy_uncertainty': 'Policy uncertainty',\n}\ncriterion_inputs = {key: W.Dropdown(options=[('Not assessed',''),('Yes','1'),('No','0'),('Not applicable','NA')], description=label+':', style={'description_width':'190px'}, layout=W.Layout(width='450px')) for key, label in criteria.items()}\ncase_view = W.Output()\nstatus = W.Output()\nsave_button = W.Button(description='Save my review', button_style='success')\n\ndef show_case(*_):\n data = pd.read_csv(review_path, keep_default_na=False)\n row = data.loc[data['case_id'] == case_choice.value].iloc[0]\n score_choice.value = str(row['overall_quality_0_to_3'])\n note_input.value = str(row['failure_description_in_my_own_words'])\n for key, field in criterion_inputs.items():\n field.value = str(row[key])\n with case_view:\n clear_output()\n display(HTML(f\"<h3>{escape(row['case_id'])}: {escape(row['platform'])} / {escape(row['event'])}</h3><b>Prompt</b><pre style='white-space:pre-wrap;overflow-wrap:anywhere;font:inherit'>{escape(str(row['prompt']))}</pre><b>Model response</b><pre style='white-space:pre-wrap;overflow-wrap:anywhere;font:inherit'>{escape(str(row['response']))}</pre>\"))\n\n\ndef save_review(_):\n with status:\n clear_output()\n if score_choice.value == '' or not note_input.value.strip():\n print('Choose a score and write your own observation before saving.')\n return\n data = pd.read_csv(review_path, keep_default_na=False)\n match = data['case_id'] == case_choice.value\n data.loc[match, 'overall_quality_0_to_3'] = score_choice.value\n data.loc[match, 'failure_description_in_my_own_words'] = note_input.value.strip()\n for key, field in criterion_inputs.items():\n data.loc[match, key] = field.value\n data.to_csv(review_path, index=False)\n complete = ((data['overall_quality_0_to_3'].astype(str) != '') & (data['failure_description_in_my_own_words'].astype(str) != '')).sum()\n print(f'Saved {case_choice.value}. Reviewed {complete} of {len(data)} cases. Download the updated CSV after reviewing.')\n\ncase_choice.observe(show_case, names='value')\nsave_button.on_click(save_review)\ndisplay(W.VBox([W.HTML('<b>Read the whole response and decide for yourself. Score 0–3; the six checks are optional.</b>'), case_choice, case_view, score_choice, note_input, *criterion_inputs.values(), save_button, status]))\nshow_case()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T11:19:30.523344Z","iopub.execute_input":"2026-09-23T11:19:30.524076Z","iopub.status.idle":"2026-09-23T11:19:30.580073Z","shell.execute_reply.started":"2026-09-23T11:19:30.524042Z","shell.execute_reply":"2026-09-23T11:19:30.579267Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 8. Summarise only the applicant's completed labels; never infer missing judgments\nimport pandas as pd\nfrom IPython.display import display\nfrom pathlib import Path\n\nreview_path = Path('/kaggle/working/cloud_runtime_eval/human_review_full_FILL_THIS_YOURSELF.csv')\ndata = pd.read_csv(review_path, keep_default_na=False)\nassert len(data) == 24 and data['case_id'].is_unique, 'Expected 24 unique cases.'\nreviewed = data[(data['overall_quality_0_to_3'].astype(str) != '') & (data['failure_description_in_my_own_words'].astype(str).str.strip() != '')].copy()\nprint(f'Your reviews completed: {len(reviewed)} / {len(data)}')\nif len(reviewed) < len(data):\n print('Still to review:', ', '.join(data.loc[~data['case_id'].isin(reviewed['case_id']), 'case_id']))\n print('Do not draw overall conclusions yet; return after all 24 are reviewed.')\nelse:\n reviewed['score'] = pd.to_numeric(reviewed['overall_quality_0_to_3'], errors='raise')\n assert reviewed['score'].between(0, 3).all(), 'Every score must be 0, 1, 2, or 3.'\n for grouping in ['platform', 'event']:\n print(f'By {grouping}:')\n table = reviewed.groupby(grouping).agg(\n cases=('score', 'size'),\n mean_score=('score', 'mean'),\n scores_0_or_1=('score', lambda values: int((values <= 1).sum())),\n ).reset_index()\n table['mean_score'] = table['mean_score'].round(2)\n display(table)\n table.to_csv(review_path.parent / f'summary_by_{grouping}.csv', index=False)\n print('Read your written failure descriptions to identify recurring patterns; the tables alone do not explain why the model failed.')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-23T11:21:08.240953Z","iopub.execute_input":"2026-09-23T11:21:08.241654Z","iopub.status.idle":"2026-09-23T11:21:08.262455Z","shell.execute_reply.started":"2026-09-23T11:21:08.241621Z","shell.execute_reply":"2026-09-23T11:21:08.261707Z"}},"outputs":[],"execution_count":null}]} |
Xet Storage Details
- Size:
- 23.1 kB
- Xet hash:
- ad28e85868377f0fdbf2bbbd9883dbac885e288c4445c83d21861bb4ea9812c3
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.