theDocWho commited on
Commit
229672c
Β·
1 Parent(s): 32401b7

Colab notebooks: use Kaggle KAGGLE_USERNAME/KAGGLE_KEY env vars

Browse files

Kaggle now hands out an API token (username + token string) instead of a
kaggle.json file. Both Colab notebooks now resolve creds in this order:

1. KAGGLE_USERNAME + KAGGLE_KEY already in os.environ (e.g. Colab Secrets).
2. Drive token file at MyDrive/ccdp/kaggle_token.txt (KEY=value lines).
3. Legacy kaggle.json (fallback for users with the old token).
4. Interactive getpass prompt, with an offer to persist to (2) for next session.

The Kaggle CLI/SDK reads the env vars directly so no kaggle.json is needed.

notebooks/colab_train_identifier.ipynb CHANGED
@@ -4,29 +4,7 @@
4
  "cell_type": "markdown",
5
  "id": "e0a4182e",
6
  "metadata": {},
7
- "source": [
8
- "# ccdp β€” train the VMMRdb identifier on Colab (GDrive-backed, crash-resumable)\n",
9
- "\n",
10
- "Same training loop as the Kaggle notebook (`ccdp train identifier-continue`), but built\n",
11
- "for Google Colab. Checkpoints write straight to Google Drive every epoch, so a runtime\n",
12
- "disconnect at epoch 9/12 doesn't lose work β€” the next launch auto-resumes from `last.pt`.\n",
13
- "\n",
14
- "### Before you run\n",
15
- "1. **Runtime β†’ Change runtime type β†’ GPU** (T4 free tier or A100 if you have Pro).\n",
16
- "2. You'll be prompted to mount your Drive in section 2. Approve it.\n",
17
- "3. You'll need a `kaggle.json` API token to pull the VMMRdb dataset\n",
18
- " (Kaggle β†’ Account β†’ Create New Token). Either:\n",
19
- " - drop it into `MyDrive/kaggle/kaggle.json` once, or\n",
20
- " - upload it via the file picker each session.\n",
21
- "\n",
22
- "### Why GDrive?\n",
23
- "Colab's `/content` is ephemeral β€” it dies with the VM. `/content/drive/MyDrive/...` is durable.\n",
24
- "We point the run dir at GDrive so every `epoch_NNN.pt` lands on Drive immediately.\n",
25
- "\n",
26
- "### Time + budget\n",
27
- "~12 epochs on a T4 β‰ˆ **6–10 h**. Free Colab sessions cap at ~12 h with idle disconnects;\n",
28
- "Colab Pro avoids most of that. Either way, the resume logic below makes a mid-run kill cheap.\n"
29
- ]
30
  },
31
  {
32
  "cell_type": "markdown",
@@ -109,12 +87,7 @@
109
  "cell_type": "markdown",
110
  "id": "1b7431c4",
111
  "metadata": {},
112
- "source": [
113
- "## 4. Fetch VMMRdb via Kaggle API\n",
114
- "\n",
115
- "The CC0 mirror is `prabashwara/vmmrdb-dataset` (~291k images, ~9,170 classes). One-time:\n",
116
- "upload your `kaggle.json` to Drive (`MyDrive/kaggle/kaggle.json`)."
117
- ]
118
  },
119
  {
120
  "cell_type": "code",
@@ -122,39 +95,7 @@
122
  "id": "79c6621b",
123
  "metadata": {},
124
  "outputs": [],
125
- "source": [
126
- "import os, shutil, pathlib\n",
127
- "KAGGLE_JSON_DRIVE = '/content/drive/MyDrive/kaggle/kaggle.json'\n",
128
- "kdir = pathlib.Path.home() / '.kaggle'\n",
129
- "kdir.mkdir(exist_ok=True)\n",
130
- "if not (kdir / 'kaggle.json').exists():\n",
131
- " assert os.path.exists(KAGGLE_JSON_DRIVE), (\n",
132
- " f'put your kaggle API token at {KAGGLE_JSON_DRIVE} (or upload via Files panel)')\n",
133
- " shutil.copy(KAGGLE_JSON_DRIVE, kdir / 'kaggle.json')\n",
134
- " os.chmod(kdir / 'kaggle.json', 0o600)\n",
135
- "!pip -q install kaggle\n",
136
- "\n",
137
- "# Cache the dataset on Drive too β€” re-downloading 70GB every session is brutal.\n",
138
- "DATA_ROOT_DRIVE = f'{DRIVE_ROOT}/data/raw/vmmrdb-dataset/prabashwara/vmmrdb-dataset'\n",
139
- "os.makedirs(DATA_ROOT_DRIVE, exist_ok=True)\n",
140
- "# Mirror to the loader's expected path:\n",
141
- "os.makedirs('data/raw/vmmrdb-dataset/prabashwara', exist_ok=True)\n",
142
- "local_data = pathlib.Path('data/raw/vmmrdb-dataset/prabashwara/vmmrdb-dataset')\n",
143
- "if local_data.exists() or local_data.is_symlink():\n",
144
- " local_data.unlink()\n",
145
- "local_data.symlink_to(DATA_ROOT_DRIVE)\n",
146
- "\n",
147
- "# Only download if Drive copy is empty.\n",
148
- "import glob\n",
149
- "if not glob.glob(f'{DATA_ROOT_DRIVE}/*'):\n",
150
- " !kaggle datasets download -d prabashwara/vmmrdb-dataset -p {DATA_ROOT_DRIVE} --unzip\n",
151
- "else:\n",
152
- " print('VMMRdb already cached on Drive β€” skipping download')\n",
153
- "\n",
154
- "from ccdp.data import vmmrdb\n",
155
- "counts = vmmrdb._class_dir_counts('data/raw/vmmrdb-dataset')\n",
156
- "print(f'class folders: {len(counts)} total images: {sum(counts.values())}')"
157
- ]
158
  },
159
  {
160
  "cell_type": "markdown",
@@ -293,4 +234,4 @@
293
  },
294
  "nbformat": 4,
295
  "nbformat_minor": 5
296
- }
 
4
  "cell_type": "markdown",
5
  "id": "e0a4182e",
6
  "metadata": {},
7
+ "source": "# ccdp β€” train the VMMRdb identifier on Colab (GDrive-backed, crash-resumable)\n\nSame training loop as the Kaggle notebook (`ccdp train identifier-continue`), but built\nfor Google Colab. Checkpoints write straight to Google Drive every epoch, so a runtime\ndisconnect at epoch 9/12 doesn't lose work β€” the next launch auto-resumes from `last.pt`.\n\n### Before you run\n1. **Runtime β†’ Change runtime type β†’ GPU** (T4 free tier or A100 if you have Pro).\n2. You'll be prompted to mount your Drive in section 2. Approve it.\n3. **Kaggle API token** for the VMMRdb dataset. Kaggle β†’ Settings β†’ **API β†’ Create New Token**.\n The new flow gives you a **username + token** pair (not a JSON file). You'll paste them\n into a `getpass` prompt in section 4. They're held in env vars `KAGGLE_USERNAME` /\n `KAGGLE_KEY` for the session only β€” nothing is written to Drive unless you ask.\n - Optional: stash them once in `MyDrive/ccdp/kaggle_token.txt` (one line per var:\n `KAGGLE_USERNAME=…`, `KAGGLE_KEY=…`) and the cell will auto-load them on later runs.\n - Legacy `kaggle.json` is still supported as a fallback.\n\n### Why GDrive?\nColab's `/content` is ephemeral β€” it dies with the VM. `/content/drive/MyDrive/...` is durable.\nWe point the run dir at GDrive so every `epoch_NNN.pt` lands on Drive immediately.\n\n### Time + budget\n~12 epochs on a T4 β‰ˆ **6–10 h**. Free Colab sessions cap at ~12 h with idle disconnects;\nColab Pro avoids most of that. Either way, the resume logic below makes a mid-run kill cheap.\n"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8
  },
9
  {
10
  "cell_type": "markdown",
 
87
  "cell_type": "markdown",
88
  "id": "1b7431c4",
89
  "metadata": {},
90
+ "source": "## 4. Fetch VMMRdb via Kaggle API\n\nThe CC0 mirror is `prabashwara/vmmrdb-dataset` (~291k images, ~9,170 classes).\n\n**Auth resolution order** β€” the cell tries each in turn and stops at the first that works:\n1. `KAGGLE_USERNAME` / `KAGGLE_KEY` already in `os.environ` (e.g. set via Colab Secrets).\n2. `MyDrive/ccdp/kaggle_token.txt` with two `KEY=value` lines.\n3. Legacy `MyDrive/kaggle/kaggle.json`.\n4. Interactive `getpass` prompt (paste username + token from Kaggle β†’ Settings β†’ API).\n"
 
 
 
 
 
91
  },
92
  {
93
  "cell_type": "code",
 
95
  "id": "79c6621b",
96
  "metadata": {},
97
  "outputs": [],
98
+ "source": "import os, json, pathlib, glob, getpass\n\nDRIVE_TOKEN_FILE = '/content/drive/MyDrive/ccdp/kaggle_token.txt'\nDRIVE_LEGACY_JSON = '/content/drive/MyDrive/kaggle/kaggle.json'\n\ndef _resolve_kaggle_creds():\n # 1. already in env (e.g. Colab Secrets injected)\n if os.environ.get('KAGGLE_USERNAME') and os.environ.get('KAGGLE_KEY'):\n return 'env'\n # 2. token file on Drive (KAGGLE_USERNAME=… / KAGGLE_KEY=…)\n if os.path.exists(DRIVE_TOKEN_FILE):\n for line in pathlib.Path(DRIVE_TOKEN_FILE).read_text().splitlines():\n line = line.strip()\n if not line or line.startswith('#') or '=' not in line:\n continue\n k, v = line.split('=', 1)\n os.environ[k.strip()] = v.strip().strip('\"').strip(\"'\")\n if os.environ.get('KAGGLE_USERNAME') and os.environ.get('KAGGLE_KEY'):\n return f'drive token file ({DRIVE_TOKEN_FILE})'\n # 3. legacy kaggle.json on Drive\n if os.path.exists(DRIVE_LEGACY_JSON):\n data = json.loads(pathlib.Path(DRIVE_LEGACY_JSON).read_text())\n os.environ['KAGGLE_USERNAME'] = data['username']\n os.environ['KAGGLE_KEY'] = data['key']\n return f'legacy kaggle.json ({DRIVE_LEGACY_JSON})'\n # 4. interactive prompt\n print('No Kaggle creds found. Paste them from Kaggle β†’ Settings β†’ API.')\n os.environ['KAGGLE_USERNAME'] = input('KAGGLE_USERNAME: ').strip()\n os.environ['KAGGLE_KEY'] = getpass.getpass('KAGGLE_KEY (token): ').strip()\n # offer to stash them on Drive for next session\n save = input(f'Save to {DRIVE_TOKEN_FILE} for next session? [y/N] ').strip().lower()\n if save == 'y':\n pathlib.Path(DRIVE_TOKEN_FILE).parent.mkdir(parents=True, exist_ok=True)\n pathlib.Path(DRIVE_TOKEN_FILE).write_text(\n f\"KAGGLE_USERNAME={os.environ['KAGGLE_USERNAME']}\\n\"\n f\"KAGGLE_KEY={os.environ['KAGGLE_KEY']}\\n\"\n )\n os.chmod(DRIVE_TOKEN_FILE, 0o600)\n print(f'saved β†’ {DRIVE_TOKEN_FILE}')\n return 'interactive'\n\nsource = _resolve_kaggle_creds()\nprint(f'[kaggle] auth via {source} (user={os.environ.get(\"KAGGLE_USERNAME\")})')\n\n# The Kaggle CLI/SDK reads KAGGLE_USERNAME/KAGGLE_KEY directly β€” no kaggle.json needed.\n!pip -q install kaggle\n\n# Cache the dataset on Drive too β€” re-downloading 70GB every session is brutal.\nDATA_ROOT_DRIVE = f'{DRIVE_ROOT}/data/raw/vmmrdb-dataset/prabashwara/vmmrdb-dataset'\nos.makedirs(DATA_ROOT_DRIVE, exist_ok=True)\n# Mirror to the loader's expected path:\nos.makedirs('data/raw/vmmrdb-dataset/prabashwara', exist_ok=True)\nlocal_data = pathlib.Path('data/raw/vmmrdb-dataset/prabashwara/vmmrdb-dataset')\nif local_data.exists() or local_data.is_symlink():\n local_data.unlink()\nlocal_data.symlink_to(DATA_ROOT_DRIVE)\n\n# Only download if Drive copy is empty.\nif not glob.glob(f'{DATA_ROOT_DRIVE}/*'):\n !kaggle datasets download -d prabashwara/vmmrdb-dataset -p {DATA_ROOT_DRIVE} --unzip\nelse:\n print('VMMRdb already cached on Drive β€” skipping download')\n\nfrom ccdp.data import vmmrdb\ncounts = vmmrdb._class_dir_counts('data/raw/vmmrdb-dataset')\nprint(f'class folders: {len(counts)} total images: {sum(counts.values())}')"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
99
  },
100
  {
101
  "cell_type": "markdown",
 
234
  },
235
  "nbformat": 4,
236
  "nbformat_minor": 5
237
+ }
notebooks/train_damage_seg_hitl_pathA.ipynb CHANGED
@@ -203,32 +203,7 @@
203
  "id": "de6ed90c",
204
  "metadata": {},
205
  "outputs": [],
206
- "source": [
207
- "import os, glob\n",
208
- "HITL_SLUG = 'YOUR_USERNAME/hitl-vehicle-damage' # <-- update with actual Kaggle slug\n",
209
- "HITL_DRIVE_CACHE = f'{DURABLE_ROOT}/data/raw/hitl'\n",
210
- "os.makedirs(HITL_DRIVE_CACHE, exist_ok=True)\n",
211
- "\n",
212
- "# Kaggle creds\n",
213
- "if PLATFORM == 'colab':\n",
214
- " import shutil, pathlib\n",
215
- " src = '/content/drive/MyDrive/kaggle/kaggle.json'\n",
216
- " kdir = pathlib.Path.home()/'.kaggle'; kdir.mkdir(exist_ok=True)\n",
217
- " if os.path.exists(src) and not (kdir/'kaggle.json').exists():\n",
218
- " shutil.copy(src, kdir/'kaggle.json'); os.chmod(kdir/'kaggle.json', 0o600)\n",
219
- "\n",
220
- "# CarDD β€” uses our packaged builder\n",
221
- "from ccdp.data import cardd_yolo\n",
222
- "cardd_yaml = cardd_yolo.build()\n",
223
- "print('CarDD yaml:', cardd_yaml)\n",
224
- "\n",
225
- "# HITL β€” download if not cached on Drive\n",
226
- "if not glob.glob(f'{HITL_DRIVE_CACHE}/*'):\n",
227
- " !pip -q install kaggle\n",
228
- " !kaggle datasets download -d {HITL_SLUG} -p {HITL_DRIVE_CACHE} --unzip\n",
229
- "else:\n",
230
- " print('HITL cache present:', HITL_DRIVE_CACHE)"
231
- ]
232
  },
233
  {
234
  "cell_type": "markdown",
@@ -524,4 +499,4 @@
524
  },
525
  "nbformat": 4,
526
  "nbformat_minor": 5
527
- }
 
203
  "id": "de6ed90c",
204
  "metadata": {},
205
  "outputs": [],
206
+ "source": "import os, json, pathlib, glob, getpass\n\nHITL_SLUG = 'YOUR_USERNAME/hitl-vehicle-damage' # <-- update with actual Kaggle slug\nHITL_DRIVE_CACHE = f'{DURABLE_ROOT}/data/raw/hitl'\nos.makedirs(HITL_DRIVE_CACHE, exist_ok=True)\n\n# --- Kaggle creds ------------------------------------------------------\n# New flow: KAGGLE_USERNAME + KAGGLE_KEY env vars (token from Kaggle β†’ Settings β†’ API).\n# Fallbacks: Drive token file -> legacy kaggle.json -> interactive prompt.\ndef _resolve_kaggle_creds():\n if os.environ.get('KAGGLE_USERNAME') and os.environ.get('KAGGLE_KEY'):\n return 'env'\n if PLATFORM == 'kaggle':\n # Kaggle kernels are auto-authenticated for their own datasets β€” nothing to do.\n return 'kaggle-native'\n # Colab: try the durable token file, then legacy json, then prompt.\n token_file = '/content/drive/MyDrive/ccdp/kaggle_token.txt'\n legacy_json = '/content/drive/MyDrive/kaggle/kaggle.json'\n if os.path.exists(token_file):\n for line in pathlib.Path(token_file).read_text().splitlines():\n line = line.strip()\n if not line or line.startswith('#') or '=' not in line: continue\n k, v = line.split('=', 1)\n os.environ[k.strip()] = v.strip().strip('\"').strip(\"'\")\n if os.environ.get('KAGGLE_USERNAME') and os.environ.get('KAGGLE_KEY'):\n return f'drive token file ({token_file})'\n if os.path.exists(legacy_json):\n data = json.loads(pathlib.Path(legacy_json).read_text())\n os.environ['KAGGLE_USERNAME'] = data['username']\n os.environ['KAGGLE_KEY'] = data['key']\n return f'legacy kaggle.json ({legacy_json})'\n print('No Kaggle creds found. Paste them from Kaggle β†’ Settings β†’ API.')\n os.environ['KAGGLE_USERNAME'] = input('KAGGLE_USERNAME: ').strip()\n os.environ['KAGGLE_KEY'] = getpass.getpass('KAGGLE_KEY (token): ').strip()\n if input(f'Save to {token_file} for next session? [y/N] ').strip().lower() == 'y':\n pathlib.Path(token_file).parent.mkdir(parents=True, exist_ok=True)\n pathlib.Path(token_file).write_text(\n f\"KAGGLE_USERNAME={os.environ['KAGGLE_USERNAME']}\\n\"\n f\"KAGGLE_KEY={os.environ['KAGGLE_KEY']}\\n\"\n )\n os.chmod(token_file, 0o600)\n print(f'saved β†’ {token_file}')\n return 'interactive'\n\nsource = _resolve_kaggle_creds()\nprint(f'[kaggle] auth via {source} (user={os.environ.get(\"KAGGLE_USERNAME\",\"<kaggle-native>\")})')\n\n# CarDD β€” uses our packaged builder\nfrom ccdp.data import cardd_yolo\ncardd_yaml = cardd_yolo.build()\nprint('CarDD yaml:', cardd_yaml)\n\n# HITL β€” download if not cached on Drive\nif not glob.glob(f'{HITL_DRIVE_CACHE}/*'):\n !pip -q install kaggle\n !kaggle datasets download -d {HITL_SLUG} -p {HITL_DRIVE_CACHE} --unzip\nelse:\n print('HITL cache present:', HITL_DRIVE_CACHE)"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
207
  },
208
  {
209
  "cell_type": "markdown",
 
499
  },
500
  "nbformat": 4,
501
  "nbformat_minor": 5
502
+ }