fishxinyu's picture
download
raw
6.75 kB
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Select Subset: Videos Containing Humans\n",
"\n",
"Filter VidGen-1M captions to retain only videos that contain humans, using keyword matching on captions."
]
},
{
"cell_type": "code",
"execution_count": 1,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Total videos: 3000\n"
]
}
],
"source": [
"import json\n",
"import re\n",
"from pathlib import Path\n",
"\n",
"CAPTION_FILE = \"/mnt/data/xinyuy/datasets/VIDGEN-1M/VidGen_1M_sample_3000.json\"\n",
"OUTPUT_FILE = \"vidgen_humans_subset.json\"\n",
"\n",
"with open(CAPTION_FILE) as f:\n",
" data = json.load(f)\n",
"\n",
"print(f\"Total videos: {len(data)}\")"
]
},
{
"cell_type": "code",
"execution_count": 2,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Videos with humans : 2540 / 3000 (84.7%)\n"
]
}
],
"source": [
"# Keywords that indicate human presence.\n",
"# Using word boundaries to avoid false positives (e.g. 'manage' matching 'man').\n",
"HUMAN_PATTERNS = [\n",
" # Generic person references\n",
" r\"\\bperson\\b\", r\"\\bpeople\\b\", r\"\\bhuman\\b\", r\"\\bhumans\\b\",\n",
" # Gender / age\n",
" r\"\\bman\\b\", r\"\\bmen\\b\", r\"\\bwoman\\b\", r\"\\bwomen\\b\",\n",
" r\"\\bboy\\b\", r\"\\bgirl\\b\", r\"\\bchild\\b\", r\"\\bchildren\\b\",\n",
" r\"\\bbaby\\b\", r\"\\bbabies\\b\", r\"\\btoddler\\b\", r\"\\bteen\\b\", r\"\\bteenager\\b\",\n",
" r\"\\bguy\\b\", r\"\\bguys\\b\", r\"\\bgentleman\\b\", r\"\\bgentlemen\\b\", r\"\\blady\\b\", r\"\\bladies\\b\",\n",
" # Pronouns used for people (capitalised too)\n",
" r\"\\bhe\\b\", r\"\\bshe\\b\", r\"\\bhis\\b\", r\"\\bher\\b\",\n",
" r\"\\bHe\\b\", r\"\\bShe\\b\", r\"\\bHis\\b\", r\"\\bHer\\b\",\n",
" # Roles / occupations (s? to catch plurals)\n",
" r\"\\bplayers?\\b\", r\"\\bathletes?\\b\", r\"\\bactors?\\b\", r\"\\bactress(es)?\\b\",\n",
" r\"\\bspeakers?\\b\", r\"\\bpresenters?\\b\", r\"\\bhosts?\\b\",\n",
" r\"\\bchefs?\\b\", r\"\\bcooks?\\b\", r\"\\bdoctors?\\b\", r\"\\bnurses?\\b\",\n",
" r\"\\bteachers?\\b\", r\"\\bstudents?\\b\", r\"\\bdrivers?\\b\", r\"\\bworkers?\\b\",\n",
" # Body parts that strongly imply a visible person\n",
" r\"\\bface\\b\", r\"\\bhands?\\b\", r\"\\bfingers?\\b\",\n",
" r\"\\bbody\\b\", r\"\\barms?\\b\", r\"\\blegs?\\b\",\n",
"]\n",
"\n",
"human_re = re.compile(\"|\".join(HUMAN_PATTERNS))\n",
"\n",
"def contains_human(caption: str) -> bool:\n",
" return bool(human_re.search(caption))\n",
"\n",
"human_videos = [item for item in data if contains_human(item[\"caption\"])]\n",
"\n",
"print(f\"Videos with humans : {len(human_videos)} / {len(data)} ({100 * len(human_videos) / len(data):.1f}%)\")"
]
},
{
"cell_type": "code",
"execution_count": 3,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"=== HUMAN (first 5) ===\n",
" FQONgrM0nB0-Scene-0020\n",
" The video shows a woman sitting in the driver's seat of a car. She is wearing a pink shirt and sunglasses. The woman is smiling and talking to the cam\n",
"\n",
" E8_7jCnTw3A-Scene-0170\n",
" The video shows a cartoon of a family of pigs having a birthday party. The mother pig is wearing a pink dress and is standing behind a table with a bi\n",
"\n",
" HlgJ13hsNR0-Scene-0060\n",
" The video shows a rectangular swimming pool with blue water. The pool is surrounded by green grass and trees. The pool is empty, and there are no peop\n",
"\n",
" LVLLHxqc9xs-Scene-0163\n",
" In the video, a woman with pink hair is seen speaking to the camera. She is wearing a black shirt and has a neutral expression on her face. The backgr\n",
"\n",
" Zt1VcESZIIw-Scene-0043\n",
" The video shows a close-up of a person's hands as they cut a piece of meat. The person is using a knife to cut the meat, which appears to be a large c\n",
"\n",
"=== NON-HUMAN (first 5) ===\n",
" J_e0mJ0AqQU-Scene-0229\n",
" The video shows a close-up of two glasses of water sitting on a car seat. The glasses are clear and filled with water up to the same level. The water \n",
"\n",
" iWADvm6DLn0-Scene-0011\n",
" The video shows a rocky trail in a forest. The trail is made up of large rocks and boulders, and there are trees on either side of the trail. The leav\n",
"\n",
" bkVRid-GbKs-Scene-0287\n",
" The video shows a cartoon shark lifting weights in a gym. The shark is seen lifting a barbell with weights on it, and then flexing its muscles. The gy\n",
"\n",
" dykexNaiYZE-Scene-0012\n",
" The video shows a close-up of a music sheet with Korean characters written on it. The camera pans across the sheet, showing the musical notes and symb\n",
"\n",
" -cXO1xtu7xw-Scene-0016\n",
" The video shows a close-up of a mechanical device with a circular component that has a green substance on it. The substance appears to be some kind of\n",
"\n"
]
}
],
"source": [
"# Spot-check a few positives and negatives\n",
"print(\"=== HUMAN (first 5) ===\")\n",
"for item in human_videos[:5]:\n",
" print(f\" {item['vid']}\")\n",
" print(f\" {item['caption'][:150]}\")\n",
" print()\n",
"\n",
"non_human = [item for item in data if not contains_human(item[\"caption\"])]\n",
"print(\"=== NON-HUMAN (first 5) ===\")\n",
"for item in non_human[:5]:\n",
" print(f\" {item['vid']}\")\n",
" print(f\" {item['caption'][:150]}\")\n",
" print()"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"with open(OUTPUT_FILE, \"w\") as f:\n",
" json.dump(human_videos, f, indent=2)\n",
"\n",
"print(f\"Saved {len(human_videos)} entries to {OUTPUT_FILE}\")"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "flexvideo",
"language": "python",
"name": "python3"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.10.20"
}
},
"nbformat": 4,
"nbformat_minor": 5
}

Xet Storage Details

Size:
6.75 kB
·
Xet hash:
2ad754fe1520442f3ca8080945104fac90716e1ca5f0e7182072deba447e02a3

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.