YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Phase 2 β Semantic Normalization / Identity Extraction
Compliments Reference DB Pipeline
Version: 2.0.0 Date: 2026-07-30 Status: PASS
Purpose
Phase 2 reads the validated Phase 1 output and performs semantic normalization (brand β brand_norm + product_line) and identity attribute extraction from the original title. Phase 2 is purely semantic β it does NOT duplicate basic string cleaning now handled by Phase 1.
Input
| Property | Value |
|---|---|
| Source | phase1/outputs/phase1_output.parquet |
| Row Count | 4,440 |
| Column Count | 20 |
What Phase 2 Does
1. Brand Semantic Normalization
- Uses
brand_cleanfrom Phase 1 (already whitespace/case normalized) - Maps 8 cleaned brand variants to
brand_norm(Compliments/Sensations) +product_line(Core/Organic/Balance/Naturally Simple/Green/Little Ones/Sensations) - Does NOT duplicate whitespace/case cleaning (Phase 1 handles this)
2. Identity Attribute Extraction (from ORIGINAL title)
- 9 boolean flags from ORIGINAL title: is_organic, is_gluten_free, is_naturally_simple, is_sugar_free, is_unsalted, is_lactose_free, is_peanut_free, is_plant_based, is_reduced_sodium
- Fat-level extraction: fat_free / reduced_fat / regular
- Fat-percentage extraction: Numeric % from title
- Flavour extraction: 32 keyword list
- Formulation extraction: 14 keyword list
3. Title Processing
- Core-title extraction: Strip brand prefix + size + bracket content (uses
title_cleanfrom Phase 1) - Variant attributes parsing: Size string β structured JSON
4. Identity Hash
- Deterministic 14-field hash (for reference; grouping in Phase 3)
Phase 2 Does NOT
- Perform product grouping
- Perform food/non-food classification
- Use taxonomy
- Use nutrition data
- Use external data sources
- Modify original
title,upc,external_id, or other source identifiers - Duplicate string cleaning already done in Phase 1
Output Schema
The Phase 2 output (phase2_output.parquet) contains 38 columns:
Original Columns (14, preserved from Phase 1)
| # | Column | Type |
|---|---|---|
| 1 | upc | str |
| 2 | external_id | str |
| 3 | brand | str |
| 4 | title | str |
| 5 | price | float64 |
| 6 | price_currency | str |
| 7 | size | str |
| 8 | size_amount | float64 |
| 9 | size_unit | str |
| 10 | size_qty | int64 |
| 11 | size_unit_norm | str |
| 12 | image_url | str |
| 13 | source | str |
| 14 | source_url | str |
Phase 1 Cleaned Columns (6, preserved from Phase 1)
| # | Column | Type | Description |
|---|---|---|---|
| 15 | brand_clean | str | Cleaned brand (from Phase 1) |
| 16 | brand_raw | str | Original brand (from Phase 1) |
| 17 | title_clean | str | Cleaned title (from Phase 1) |
| 18 | title_raw | str | Original title (from Phase 1) |
| 19 | upc_raw | str | Original UPC (from Phase 1) |
| 20 | external_id_raw | str | Original external_id (from Phase 1) |
New Columns (18, added by Phase 2)
| # | Column | Type | Description |
|---|---|---|---|
| 21 | brand_norm | str | Normalized brand (Compliments/Sensations) |
| 22 | product_line | str | Product line (Core/Organic/Balance/Naturally Simple/Green/Little Ones/Sensations) |
| 23 | is_organic | bool | Organic flag |
| 24 | is_gluten_free | bool | Gluten-free flag |
| 25 | is_naturally_simple | bool | Naturally Simple flag |
| 26 | is_sugar_free | bool | Sugar-free flag |
| 27 | is_unsalted | bool | Unsalted flag |
| 28 | is_lactose_free | bool | Lactose-free flag |
| 29 | is_peanut_free | bool | Peanut-free flag |
| 30 | is_plant_based | bool | Plant-based flag |
| 31 | is_reduced_sodium | bool | Reduced sodium flag |
| 32 | fat_level | str | Fat level (fat_free/reduced_fat/regular) |
| 33 | fat_percentage | float64 | Fat percentage from title (nullable) |
| 34 | flavour | list | Extracted flavour keywords |
| 35 | formulation | list | Extracted formulation keywords |
| 36 | variant_attributes | str | Parsed size as JSON |
| 37 | core_title | str | Title with brand prefix + size stripped |
| 38 | identity_hash | str | Deterministic 14-field identity hash |
Validation Results
| Check | Result |
|---|---|
| Row count preserved (4,440) | PASS |
| All original columns preserved | PASS |
| All 18 new columns added | PASS |
| Brand normalization (2 values) | PASS |
| Product line (7 values) | PASS |
| Fat level values valid | PASS |
| Identity flags are boolean | PASS |
| Core titles not empty | PASS |
| Identity hashes not empty | PASS |
| No duplicate external_ids | PASS |
| Overall | PASS |
Files
phase2/
README.md
src/phase2.py
notebooks/phase2.ipynb
outputs/
phase2_output.parquet
phase2_executed.ipynb
validation/
phase2_validation.json
statistics/
phase2_statistics.json
How to Run
# Using the Python script
cd /home/sara/gsoc/compliments-reference-db
python -m phase2.src.phase2
# Using the notebook
jupyter notebook phase2/notebooks/phase2.ipynb
Handoff to Phase 3
Phase 3 must read:
phase2/outputs/phase2_output.parquet
Phase 3 must NOT:
- Re-download from HuggingFace
- Read the original products.parquet directly
- Read Phase 1 output directly
- Use any other data source
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support