wandb: Detected [huggingface_hub.inference, openai] in use. wandb: Use W&B Weave for improved LLM call tracing. Install Weave with `pip install weave` then add `import weave` to the top of your script. wandb: For more information, check out the docs at: https://weave-docs.wandb.ai/ Initial GPU Memory Stats: Allocated: 0.00 MB Cached: 0.00 MB Max Allocated: 0.00 MB Initializing bidirectional cross-attention model abhinand/MedEmbed-large-v0.1... ============================================================ đŸĻĨ UNSLOTH MODE: Loading encoder with optimized kernels ============================================================ đŸĻĨ Unsloth Status: Available: ✅ Yes FastSentenceTransformer: ✅ Yes (preferred) Version: 2026.1.4 CUDA: ✅ Yes GPU: NVIDIA GeForce RTX 4080 SUPER GPU Memory: 16.71 GB API: FastSentenceTransformer (specialized for embeddings) ✅ Target modules: ['q_proj', 'k_proj', 'v_proj', 'o_proj', 'gate_proj', 'up_proj', 'down_proj'] [INFO] Auto-detected max_seq_length for abhinand/MedEmbed-large-v0.1: 512 ============================================================ đŸĻĨ UNSLOTH MODE: Creating optimized sentence encoder ============================================================ API: FastSentenceTransformer (specialized for embeddings) Model: abhinand/MedEmbed-large-v0.1 Max seq length: 512 Full finetuning: False Unsloth: Using fast encoder path for bert (torch.compile + SDPA) `torch_dtype` is deprecated! Use `dtype` instead! Unsloth: Enabled gradient checkpointing âš ī¸ Unsloth loading failed: Target modules {'down_proj', 'v_proj', 'up_proj', 'k_proj', 'o_proj', 'q_proj', 'gate_proj'} not found in the base model. Please check the target modules and try again. Traceback (most recent call last): File "/home/dtim/Ataur/LOKI/Cross_Attention/run_cross_attention.py", line 847, in main sentence_encoder = create_unsloth_sentence_encoder( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/dtim/Ataur/LOKI/Cross_Attention/unsloth_encoder.py", line 462, in create_unsloth_sentence_encoder sentence_encoder = FastSentenceTransformer.get_peft_model( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/unsloth/models/sentence_transformer.py", line 1685, in get_peft_model peft_model = peft_get_peft_model(inner_model, lora_config) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/mapping_func.py", line 122, in get_peft_model return MODEL_TYPE_TO_PEFT_MODEL_MAPPING[peft_config.task_type]( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/peft_model.py", line 2966, in __init__ super().__init__(model, peft_config, adapter_name, **kwargs) File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/peft_model.py", line 129, in __init__ self.base_model = cls(model, {adapter_name: peft_config}, adapter_name) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/tuners/tuners_utils.py", line 298, in __init__ self.inject_adapter(self.model, adapter_name, low_cpu_mem_usage=low_cpu_mem_usage, state_dict=state_dict) File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/tuners/tuners_utils.py", line 880, in inject_adapter raise ValueError(error_msg) ValueError: Target modules {'down_proj', 'v_proj', 'up_proj', 'k_proj', 'o_proj', 'q_proj', 'gate_proj'} not found in the base model. Please check the target modules and try again. Falling back to standard SentenceTransformer loading... Encoder fine-tuning enabled (trainable_encoder=True). No encoder LoRA adapters are attached by default. Initialized model without Flash Attention Embedding dimension: 1024 Max sequence length: 512 Loading datasets... Loaded 19516 training examples ⚡ Sampled 10000 training examples (deterministic, seed=42) Loaded 2444 evaluation examples ⚡ Sampled 1000 evaluation examples (deterministic) Using top_k_sparse attention mechanism Initializing Top-K Sparse Attention (k=5) Initializing top-k sparse attention with method: zeros Successfully applied zeros initialization to top-k sparse attention Initializing top-k sparse attention with method: zeros Successfully applied zeros initialization to top-k sparse attention Using cosine pair scoring method Bidirectional model initialized with top_k=5, pair_score_method=cosine, share_weights=True, use_refinement=False Using initialization method: zeros Initialization parameters: {'bias_value': 0.0} Sentence encoder is trainable Sentence encoder dtype detected: torch.float32 Model initialized with standard cross-attention layers 📊 Model Parameter Statistics: Total parameters: 354,040,343 Trainable parameters: 354,040,343 Frozen parameters: 0 Trainable percentage: 100.00% đŸŽ¯ Loss and Aggregation Configuration: Architecture: Bidirectional Loss type: bidirectional_triplet Aggregation method: top_k_pairs Top-k value: 5 Norm type: rmsnorm Q/K RMSNorm: False Embedding caching: ❌ Disabled Triplet batching: 🔒 Isolated examples Triplet strategy: LIMITED Max triplets/example: 2 🎲 Initialization Configuration: Method: zeros Description: Initialize all weights to zero (tests uniform attention baseline) Parameters: {'bias_value': 0.0} 📊 Loss Component Weights (will be normalized): Triplet weight: 0.5 Attention weight: 0.0 Pair weight: 0.3 âš™ī¸ Loss Parameters: Margin: 0.3 Scale: 10.0 Pair margin: 0.3 Pair score method: cosine Share attention weights: True Join path extraction: True Join path threshold: 0.15 Self-attention enabled: False Attention mechanism: top_k_sparse Sparse top-k: 5 đŸŽ¯ Attention Mechanism Configuration: Attention type: top_k_sparse Sparse top-k: 5 Initialization method: zeros 🔄 Bidirectional-Specific Configuration: Pair margin: 0.3 Pair score method: cosine Share attention weights: True Join path extraction: True Join path threshold: 0.15 Self-attention enabled: False Output will be saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945 Training configuration saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_config.json Starting Model Training... ================================================== Starting training with ID-based triplet batches... đŸŽ¯ Training curves tracker initialized: 📁 Output directory: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945 📊 Tracking batch losses: True 📈 Tracking validation loss: True 🔍 Tracking row-sentence metrics: True 💾 Auto-save: True 🎨 Auto-plot: False đŸŽ¯ Training curves tracking enabled 🏆 Test-based model saving enabled - will save best models by test metrics 🔍 Loading row-sentence evaluation data... 📊 Row-sentence evaluation data: Test examples: 36 Annotations: 0 Examples with annotations: 0 â„šī¸ Row-sentence data in MIMIC format - will sync with MIMIC loader 🔍 Loading MIMIC evaluation data... [INFO] Loaded 18 MIMIC annotations Diagnosis row groundings: 114 Medication row groundings: 162 [INFO] MIMIC evaluation data loaded: Test examples: 36 Annotations: 18 admissions Matched examples: 36 ✅ MIMIC evaluation enabled (all 36 test examples) 🔄 Cache disabled: MIMIC test will be re-encoded each epoch 🔗 Syncing generic row-sentence evaluation with loaded MIMIC data for consistent AP tracking... Preparing triplet batches... Triplet sampling strategy: LIMITED Max triplets per example: 2 Using ISOLATED triplet batching (triplets from the same example stay together) Processing examples (isolated batches): 100%|██████████| 10000/10000 [00:01<00:00, 5335.99it/s] 📊 Triplet Generation Statistics (Isolated Batching): Total examples: 10000 Total triplets: 16502 Avg triplets/example: 1.65 Min triplets/example: 1 Max triplets/example: 2 Examples with 0 triplets: 0 Created 10000 triplet batches for training (all triplets included, adaptive batching) Scheduler configured for 200000 steps with 20000 warmup steps đŸ”Ĩ STAGE 0: FROZEN ENCODER ONLY EVALUATION... Evaluating with ONLY the frozen sentence encoder (no cross-attention at all)... đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:05<00:00, 4.21s/it] đŸ”Ĩ FROZEN ENCODER ONLY Accuracy: 0.806 (Total comparisons: 24296) 🔍 Stage 0: Evaluating row-sentence alignment... â„šī¸ Using MIMIC evaluator (Level A)... [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.185 Medication: 18 examples, F1=0.242 đŸ”Ĩ Stage 0 Row-Sent Avg Precision: 0.419 đŸ”Ĩ Stage 0 Row-Sent Overall Accuracy: 0.213 🚀 STAGE 1: SOPHISTICATED MODEL EVALUATION (PRE-TRAINING)... Evaluating sophisticated model BEFORE training... 🚀 Using CUDA device for evaluation đŸŽ¯ TRAINING CONFIGURATION: Starting from Stage 0 📋 Configuration detected: Encoder trainable: True Use cache: False Use LoRA: False Encoder-only training: True Mode: Encoder-only training (cross-attention heads frozen) Encoder tuning mode: gradual unfreezing 🔓 Gradual unfreezing initialized: 2/24 top layers trainable (layers 22..23) Enabling gradient checkpointing for memory efficiency Optimizer groups configured: head=0 @ 1.00e-04, encoder=391 @ 1.00e-05 Initialized EncoderOnlyTripletLoss with margin=0.3, scale=10.0, ranking=infonce Initialized EncoderOnlyTripletLoss with margin=0.3, scale=10.0 Using cosine LR schedule with warmup 🔧 Using bidirectional evaluation for bidirectional model Evaluating bidirectional model with join path extraction... Evaluating examples: 100%|██████████| 1000/1000 [1:10:13<00:00, 4.21s/it] 🚀 SOPHISTICATED MODEL Accuracy (Pre-training): 0.806 🔍 Stage 1: Evaluating row-sentence alignment (Sophisticated Untrained)... [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.185 Medication: 18 examples, F1=0.242 🚀 Stage 1 Row-Sent Avg Precision: 0.418 🚀 Stage 1 Row-Sent Overall Accuracy: 0.213 📊 ARCHITECTURE BENEFIT: đŸ”Ĩ→🚀 Sophisticated model benefit: +0.000 (0.0%) 📈 ROW-SENTENCE PROGRESSION (Avg Precision): đŸ”Ĩ Stage 0 (Frozen): 0.419 🚀 Stage 1 (Sophisticated): 0.418 (+-0.000) âš ī¸ WARNING: Sophisticated model barely improves over frozen encoder! Consider if the architecture complexity is necessary for this task. 📊 Logging initial stage metrics to wandb... ✅ Baseline metrics logged to wandb (epoch 0) 📊 Adding initial stage metrics to training curves... 📊 Initial stage metrics recorded: đŸ”Ĩ Stage 0 (Frozen): 0.81 acc, 0.42 row-sent AP 🚀 Stage 1 (Sophisticated Untrained): 0.81 acc, 0.42 row-sent AP đŸŽ¯ Stage 2 (Initial): 0.00 acc, 0.00 row-sent AP 📊 Adding Epoch 0 (untrained model) baseline to training curves... 🎓 STARTING TRAINING OF ENCODER ONLY (NO HEADS) Training model type: Encoder Only (No Heads) Training from Stage: 0 ==================== Epoch 1/20 ==================== Training model has 26242048 trainable parameters Epoch 1: 100%|██████████| 10000/10000 [4:48:51<00:00, 1.73s/it, batch_loss=0.687] Current learning rates: 5.00e-06 | margin=0.300 Epoch 1 loss statistics - Mean: 0.689, Min: 0.520, Max: 0.854 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:57<00:00, 4.20s/it] Epoch 1 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36305.59it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.58it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.203 Medication: 18 examples, F1=0.244 Row-Sent Overall Acc: 0.224 Row-Sent Avg Precision: 0.435 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37328.79it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.44it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.203 Medication: 18 examples, F1=0.244 MIMIC Diagnosis - P: 0.119, R: 0.847, F1: 0.203 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244 📊 Epoch 1 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21544.8s 📈 Learning Rate: 5.00e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.44 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) New best model! Accuracy: 0.806 New best test overall accuracy! 0.224 (Epoch 1) New best test average precision! 0.435 (Epoch 1) ==================== Epoch 2/20 ==================== 🔓 Gradual unfreezing: enabled layers 21..21 (total 3/24) Epoch 2: 100%|██████████| 10000/10000 [4:49:35<00:00, 1.74s/it, batch_loss=0.683] Current learning rates: 1.00e-05 | margin=0.300 Epoch 2 loss statistics - Mean: 0.689, Min: 0.523, Max: 0.871 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:21<00:00, 4.22s/it] Epoch 2 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36314.32it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.60it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.203 Medication: 18 examples, F1=0.244 Row-Sent Overall Acc: 0.223 Row-Sent Avg Precision: 0.435 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37748.74it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.77it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.203 Medication: 18 examples, F1=0.244 MIMIC Diagnosis - P: 0.118, R: 0.847, F1: 0.203 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244 📊 Epoch 2 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21613.4s 📈 Learning Rate: 1.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.44 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 3/20 ==================== 🔓 Gradual unfreezing: enabled layers 20..20 (total 4/24) Epoch 3: 100%|██████████| 10000/10000 [4:50:32<00:00, 1.74s/it, batch_loss=0.683] Current learning rates: 9.92e-06 | margin=0.300 Epoch 3 loss statistics - Mean: 0.689, Min: 0.504, Max: 0.875 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:17<00:00, 4.22s/it] Epoch 3 Accuracy: 0.805 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 35511.51it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.69it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.435 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36846.01it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.21it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.244 MIMIC Diagnosis - P: 0.117, R: 0.847, F1: 0.201 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244 📊 Epoch 3 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21666.8s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.44 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 4/20 ==================== 🔓 Gradual unfreezing: enabled layers 19..19 (total 5/24) Epoch 4: 100%|██████████| 10000/10000 [4:50:41<00:00, 1.74s/it, batch_loss=0.683] Current learning rates: 9.70e-06 | margin=0.300 Epoch 4 loss statistics - Mean: 0.689, Min: 0.525, Max: 0.854 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:11<00:00, 4.21s/it] Epoch 4 Accuracy: 0.805 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36756.32it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.76it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.244 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.435 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 18810.88it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 56.02it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.244 MIMIC Diagnosis - P: 0.117, R: 0.847, F1: 0.201 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244 📊 Epoch 4 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21669.1s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 5/20 ==================== 🔓 Gradual unfreezing: enabled layers 18..18 (total 6/24) Epoch 5: 100%|██████████| 10000/10000 [4:50:47<00:00, 1.74s/it, batch_loss=0.686] Current learning rates: 9.33e-06 | margin=0.300 Epoch 5 loss statistics - Mean: 0.689, Min: 0.535, Max: 0.854 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:56<00:00, 4.20s/it] Epoch 5 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36054.19it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.85it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.434 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 39097.60it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.19it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 MIMIC Diagnosis - P: 0.117, R: 0.849, F1: 0.201 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.243 📊 Epoch 5 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21660.5s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 6/20 ==================== Epoch 6: 100%|██████████| 10000/10000 [4:50:51<00:00, 1.75s/it, batch_loss=0.688] Current learning rates: 8.83e-06 | margin=0.300 Epoch 6 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.857 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:18<00:00, 4.22s/it] Epoch 6 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 35603.62it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.66it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.244 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.434 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36756.32it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 58.97it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.244 MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.201 MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244 📊 Epoch 6 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21686.2s 📈 Learning Rate: 8.83e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 7/20 ==================== Epoch 7: 100%|██████████| 10000/10000 [4:51:28<00:00, 1.75s/it, batch_loss=0.682] Current learning rates: 8.21e-06 | margin=0.300 Epoch 7 loss statistics - Mean: 0.689, Min: 0.512, Max: 0.866 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:17<00:00, 4.22s/it] Epoch 7 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37504.95it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.92it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.16it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.435 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 16710.37it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 57.66it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.201 MIMIC Medication - P: 0.152, R: 0.840, F1: 0.243 📊 Epoch 7 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21722.1s 📈 Learning Rate: 8.21e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 8/20 ==================== Epoch 8: 100%|██████████| 10000/10000 [4:51:15<00:00, 1.75s/it, batch_loss=0.682] Current learning rates: 7.50e-06 | margin=0.300 Epoch 8 loss statistics - Mean: 0.689, Min: 0.536, Max: 0.865 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:14<00:00, 4.21s/it] Epoch 8 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37099.49it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.28it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.201 Medication: 18 examples, F1=0.243 Row-Sent Overall Acc: 0.222 Row-Sent Avg Precision: 0.434 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36587.10it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.93it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.243 MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.200 MIMIC Medication - P: 0.152, R: 0.840, F1: 0.243 📊 Epoch 8 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21706.6s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 9/20 ==================== Epoch 9: 100%|██████████| 10000/10000 [4:51:10<00:00, 1.75s/it, batch_loss=0.681] Current learning rates: 6.71e-06 | margin=0.300 Epoch 9 loss statistics - Mean: 0.689, Min: 0.522, Max: 0.859 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:14<00:00, 4.21s/it] Epoch 9 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 33149.27it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.62it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 Row-Sent Overall Acc: 0.221 Row-Sent Avg Precision: 0.434 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 35704.65it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 58.34it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200 MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241 📊 Epoch 9 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 1) âąī¸ Epoch Time: 21701.2s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 10/20 ==================== Epoch 10: 100%|██████████| 10000/10000 [4:51:09<00:00, 1.75s/it, batch_loss=0.678] Current learning rates: 5.87e-06 | margin=0.300 Epoch 10 loss statistics - Mean: 0.689, Min: 0.531, Max: 0.860 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:17<00:00, 4.22s/it] Epoch 10 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 34735.44it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.24it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 Row-Sent Overall Acc: 0.221 Row-Sent Avg Precision: 0.434 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 38101.17it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.52it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200 MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241 📊 Epoch 10 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21704.0s 📈 Learning Rate: 5.87e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) New best model! Accuracy: 0.806 ==================== Epoch 11/20 ==================== Epoch 11: 100%|██████████| 10000/10000 [4:51:28<00:00, 1.75s/it, batch_loss=0.678] Current learning rates: 5.00e-06 | margin=0.300 Epoch 11 loss statistics - Mean: 0.689, Min: 0.530, Max: 0.859 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:21<00:00, 4.22s/it] Epoch 11 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 35229.80it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.53it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 Row-Sent Overall Acc: 0.221 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 34976.82it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.47it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.17it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.241 MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200 MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241 📊 Epoch 11 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21726.3s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 12/20 ==================== Epoch 12: 100%|██████████| 10000/10000 [4:51:23<00:00, 1.75s/it, batch_loss=0.685] Current learning rates: 4.13e-06 | margin=0.300 Epoch 12 loss statistics - Mean: 0.689, Min: 0.508, Max: 0.865 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:16<00:00, 4.22s/it] Epoch 12 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 33391.19it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.82it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.220 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37588.98it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.78it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.11it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.200 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.116, R: 0.847, F1: 0.200 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 12 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21717.2s 📈 Learning Rate: 4.13e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 13/20 ==================== Epoch 13: 100%|██████████| 10000/10000 [4:51:05<00:00, 1.75s/it, batch_loss=0.684] Current learning rates: 3.29e-06 | margin=0.300 Epoch 13 loss statistics - Mean: 0.689, Min: 0.528, Max: 0.890 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:57<00:00, 4.20s/it] Epoch 13 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36711.63it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.63it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 39808.84it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.23it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 13 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21679.7s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 14/20 ==================== Epoch 14: 100%|██████████| 10000/10000 [4:50:45<00:00, 1.74s/it, batch_loss=0.686] Current learning rates: 2.50e-06 | margin=0.300 Epoch 14 loss statistics - Mean: 0.689, Min: 0.542, Max: 0.859 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:57<00:00, 4.20s/it] Epoch 14 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36114.55it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.32it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 38796.23it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.26it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 14 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21660.1s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 15/20 ==================== Epoch 15: 100%|██████████| 10000/10000 [4:50:47<00:00, 1.74s/it, batch_loss=0.688] Current learning rates: 1.79e-06 | margin=0.300 Epoch 15 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.869 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:03<00:00, 4.20s/it] Epoch 15 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37356.49it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.65it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36428.21it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.69it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 15 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21666.9s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 16/20 ==================== Epoch 16: 100%|██████████| 10000/10000 [4:50:44<00:00, 1.74s/it, batch_loss=0.689] Current learning rates: 1.17e-06 | margin=0.300 Epoch 16 loss statistics - Mean: 0.689, Min: 0.534, Max: 0.860 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:57<00:00, 4.20s/it] Epoch 16 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 36393.09it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.83it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 39250.05it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.32it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 16 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21658.9s 📈 Learning Rate: 1.17e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 17/20 ==================== Epoch 17: 100%|██████████| 10000/10000 [4:50:45<00:00, 1.74s/it, batch_loss=0.686] Current learning rates: 1.00e-06 | margin=0.300 Epoch 17 loss statistics - Mean: 0.689, Min: 0.534, Max: 0.849 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:57<00:00, 4.20s/it] Epoch 17 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37467.73it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.80it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37035.80it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.21it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 17 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21660.2s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 18/20 ==================== Epoch 18: 100%|██████████| 10000/10000 [4:50:43<00:00, 1.74s/it, batch_loss=0.679] Current learning rates: 1.00e-06 | margin=0.300 Epoch 18 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.872 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:09:58<00:00, 4.20s/it] Epoch 18 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 37090.38it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.83it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 34647.76it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.39it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.19it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 18 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21658.0s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 19/20 ==================== Epoch 19: 100%|██████████| 10000/10000 [4:50:59<00:00, 1.75s/it, batch_loss=0.687] Current learning rates: 1.00e-06 | margin=0.300 Epoch 19 loss statistics - Mean: 0.689, Min: 0.518, Max: 0.866 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:17<00:00, 4.22s/it] Epoch 19 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 32640.50it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.64it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 34799.48it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 52.03it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 19 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21693.0s 📈 Learning Rate: 1.00e-06 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) ==================== Epoch 20/20 ==================== Epoch 20: 100%|██████████| 10000/10000 [4:51:12<00:00, 1.75s/it, batch_loss=0.684] Current learning rates: 1.00e-06 | margin=0.300 Epoch 20 loss statistics - Mean: 0.689, Min: 0.538, Max: 0.871 Loss not decreasing. Patience: 1/1 Reset learning rate to 5.00e-05 Evaluating after epoch... 🔎 Validation evaluator: Stage 0 (frozen encoder baseline) đŸ”Ĩ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs) Evaluating frozen encoder baseline: 100%|██████████| 1000/1000 [1:10:14<00:00, 4.21s/it] Epoch 20 Accuracy: 0.806 Collecting tables and contexts for test_row_sent_epoch split... Processing test_row_sent_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 34560.53it/s] Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split Using device: cuda for embedding cache Encoding test_row_sent_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 61.65it/s] Encoding test_row_sent_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.18it/s] 🔎 Row-sent evaluator: Stage 0 (frozen encoder baseline) [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 Row-Sent Overall Acc: 0.219 Row-Sent Avg Precision: 0.433 Examples evaluated: 36 Collecting tables and contexts for mimic_test_epoch split... Processing mimic_test_epoch examples: 100%|██████████| 36/36 [00:00<00:00, 38333.32it/s] Found 36 unique tables and 39 unique contexts in mimic_test_epoch split Using device: cuda for embedding cache Encoding mimic_test_epoch tables: 100%|██████████| 36/36 [00:00<00:00, 59.05it/s] Encoding mimic_test_epoch contexts: 100%|██████████| 39/39 [00:07<00:00, 5.14it/s] [INFO] MIMIC Row Grounding evaluation: 36 examples Diagnosis: 18 examples, F1=0.198 Medication: 18 examples, F1=0.240 MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198 MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240 📊 Epoch 20 Summary: đŸ”Ĩ Train Loss: 0.69 Âą 0.04 đŸŽ¯ Val Accuracy: 0.81 🏆 Best Accuracy: 0.81 (Epoch 10) âąī¸ Epoch Time: 21702.7s 📈 Learning Rate: 5.00e-05 🔍 Row-Sent Overall Acc: 0.22 🏆 Best Test Overall Acc: 0.22 (Epoch 1) 🔍 Row-Sent Avg Precision: 0.43 🏆 Best Test Avg Precision: 0.44 (Epoch 1) Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache) Loading best model from epoch 10 with accuracy 0.806 Best model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_model_epoch_10/model.pt đŸŽ¯ SAVING BEST TEST-BASED MODELS: Loading best test overall accuracy model from epoch 1 (0.224) Best test overall accuracy model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_test_overall_acc_epoch_1/model.pt Loading best test average precision model from epoch 1 (0.435) Best test average precision model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_test_avg_precision_epoch_1/model.pt Reloaded validation-based best model for final analysis ========================================================================================== đŸŽ¯ COMPREHENSIVE 3-STAGE IMPACT ANALYSIS ========================================================================================== 📊 COMPLETE PERFORMANCE BREAKDOWN: đŸ”Ĩ Stage 0 - Frozen Encoder Only: 0.806 🚀 Stage 1 - Sophisticated (Pre): 0.806 (+0.000) 🏆 Stage 2 - Trained Model: 0.806 (+0.000) 📈 Total Improvement: +0.000 (0.0%) 🏆 TEST METRICS SUMMARY: đŸŽ¯ Best Test Overall Accuracy: 0.224 (Epoch 1) đŸŽ¯ Best Test Average Precision: 0.435 (Epoch 1) 📊 Validation vs Test Performance: - Validation Best: 0.806 (Epoch 10) - Test Overall Acc: 0.224 (Epoch 1) - Test Avg Precision: 0.435 (Epoch 1) âš ī¸ NOTE: Best test overall accuracy occurred at epoch 1, not at best validation epoch 10✅ Final summary logged to wandb 🎨 Skipping 3-STAGE example visualizations (skip_four_stage_viz=True) ========================================================================================== ====================================================================== đŸŽ¯ TRAINING CURVES SUMMARY ====================================================================== ====================================================================== đŸŽ¯ TRAINING SUMMARY - abhinand/MedEmbed-large-v0.1 ====================================================================== 📊 Total Epochs: 20 🏆 Best Accuracy: 0.81 (Epoch 10) 📈 Final Accuracy: 0.81 📉 Accuracy Improvement: -0.00 đŸ”Ĩ Initial Loss: 0.69 đŸŽ¯ Final Loss: 0.69 📉 Loss Reduction: -0.00 âąī¸ Average Epoch Time: 21674.9s 🕐 Total Training Time: 433497.5s (7225.0 min) 🏆 Best Test Metrics: 🔍 Best Overall Accuracy: 0.22 (Epoch 1) 🔍 Best Average Precision: 0.44 (Epoch 1) 📊 Initial Stage Metrics: đŸ”Ĩ Stage 0 (Frozen): 0.81 acc, 0.42 row-sent AP 🚀 Stage 1 (Sophisticated Untrained): 0.81 acc, 0.42 row-sent AP đŸŽ¯ Stage 2 (Initial): 0.00 acc, 0.00 row-sent AP ====================================================================== Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_final_training_curves.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_final_training_curves.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_training_loss.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_training_loss.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_validation_accuracy.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_validation_accuracy.pdf 📈 Generating batch-level analysis... Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.pdf đŸŒĄī¸ Batch losses heatmap saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.pdf 📈 Batch losses plot for epoch 1 saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.pdf 📈 Batch losses plot for epoch 21 saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.png ✅ Training curves analysis complete! 📁 All plots saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots 💾 Training data saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_data 🎓 TRAINING COMPLETED: Returning trained Encoder Only (No Heads) ✅ Training completed! The returned model is the BEST performing model from training. 📁 Check the training logs above to see which epoch achieved the highest validation accuracy. Saved best model to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt Evaluating on Test Set... Loaded 36 test examples Cache disabled - evaluation will compute embeddings on-the-fly Evaluating bidirectional model with join path extraction... Evaluating examples: 100%|██████████| 36/36 [00:18<00:00, 1.94it/s] Saved 36 join path examples to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/extracted_join_paths.json Test Accuracy: 0.931 Join Paths Extracted: 36 Evaluation results saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/test_metrics.json 🎨 Generating visualizations using the BEST model (highest validation accuracy)... ✅ CONFIRMED: Using BEST model checkpoint from training 📁 Best model path: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt 🔧 Using aggregation method: top_k_pairs ✅ Best model file confirmed to exist Running comprehensive visualizations using top_k_pairs... 🔍 Loading trained model from: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt 📁 Loading TRAINED model from checkpoint: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt ✅ CONFIRMED: Loading BEST model checkpoint (highest validation accuracy) 📄 Found configuration file: args.json ✅ Loaded from config: model_type=Bidirectional, embedding_dim=1024, attention_type=top_k_sparse Creating dynamic encoder for 1024 dimensions on cuda 📐 Final embedding dimension: 1024 Creating BidirectionalTableTextModel... 🚀 Using top_k_sparse attention mechanism 📋 Using top-5 sparse attention Using top_k_sparse attention mechanism Initializing Top-K Sparse Attention (k=5) Initializing top-k sparse attention with method: xavier_uniform Successfully applied xavier_uniform initialization to top-k sparse attention Initializing top-k sparse attention with method: xavier_uniform Successfully applied xavier_uniform initialization to top-k sparse attention Using cosine pair scoring method Bidirectional model initialized with top_k=3, pair_score_method=cosine, share_weights=False, use_refinement=True Using initialization method: xavier_uniform Sentence encoder is frozen Sentence encoder dtype detected: torch.float32 Loading trained weights from output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt ✅ Trained weights loaded successfully (custom layers only) 🏆 VERIFICATION: Best model weights are now active for visualization đŸŽ¯ Model ready for visualization: BEST TRAINED model on cuda ✅ Trained model loaded with dynamic dimension detection on cuda Processing example 16 (ID: 2221973625275923611): 8 rows, 69 sentences Using aggregation method: top_k_pairs Generating comprehensive 4-panel analysis... 🔍 Creating comprehensive analysis for Example 16... Found 8 rows and 69 sentences Computing comprehensive similarities... Computing raw embedding similarities... Computing bidirectional attention and contextualized similarities... Applying top_k_sparse attention mechanism... Computing contextualized similarities... Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.pdf 💾 Saved comprehensive analysis to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png 📊 Analysis Summary for Example 16: Raw Embeddings - Range: [-0.101, 0.082] Contextualized Similarities - Range: [-0.071, 0.396] Final Model Similarities - Range: [-0.091, 0.082] Cross-Attention - Range: [0.000, 0.200] 🔝 Top 3 pairs by Raw Embeddings: 1. Row 2 - Sentence 57: 0.082 2. Row 6 - Sentence 55: 0.079 3. Row 6 - Sentence 45: 0.074 🔝 Top 3 pairs by Contextualized Similarities: 1. Row 1 - Sentence 3: 0.396 2. Row 3 - Sentence 2: 0.379 3. Row 5 - Sentence 4: 0.377 🔝 Top 3 pairs by Final Model Similarities: 1. Row 4 - Sentence 27: 0.082 2. Row 5 - Sentence 20: 0.082 3. Row 4 - Sentence 10: 0.079 🔝 Top 3 pairs by Cross-Attention: 1. Row 5 - Sentence 2: 0.200 2. Row 5 - Sentence 3: 0.200 3. Row 5 - Sentence 4: 0.200 📄 Detailed analysis saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/example_16_similarity_analysis.txt All visualizations saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations đŸ”Ŧ Creating comprehensive model analysis... đŸ”Ŧ Creating comprehensive analysis for example 16 using BEST model... 🔍 Creating comprehensive analysis for Example 16... Found 8 rows and 69 sentences Computing comprehensive similarities... Computing raw embedding similarities... Computing bidirectional attention and contextualized similarities... Applying top_k_sparse attention mechanism... Computing contextualized similarities... Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.pdf 💾 Saved comprehensive analysis to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png 📊 Analysis Summary for Example 16: Raw Embeddings - Range: [0.456, 0.725] Contextualized Similarities - Range: [0.834, 0.949] Final Model Similarities - Range: [0.456, 0.725] Cross-Attention - Range: [0.000, 0.200] 🔝 Top 3 pairs by Raw Embeddings: 1. Row 8 - Sentence 25: 0.725 2. Row 6 - Sentence 25: 0.723 3. Row 2 - Sentence 28: 0.717 🔝 Top 3 pairs by Contextualized Similarities: 1. Row 4 - Sentence 38: 0.949 2. Row 1 - Sentence 2: 0.943 3. Row 4 - Sentence 4: 0.943 🔝 Top 3 pairs by Final Model Similarities: 1. Row 8 - Sentence 25: 0.725 2. Row 6 - Sentence 25: 0.723 3. Row 2 - Sentence 28: 0.717 🔝 Top 3 pairs by Cross-Attention: 1. Row 5 - Sentence 2: 0.200 2. Row 5 - Sentence 3: 0.200 3. Row 5 - Sentence 4: 0.200 📄 Detailed analysis saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/example_16_similarity_analysis.txt đŸ”Ŧ Generating step-by-step diagnostics for example 16... đŸ”Ŧ Generating step-by-step diagnostics for example 16... Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step1_raw_similarities.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step1_raw_similarities.pdf Applying top_k_sparse attention mechanism... Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step2_forward_attention.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step2_forward_attention.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step3_reverse_attention.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step3_reverse_attention.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step4_contextualized_similarities.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step4_contextualized_similarities.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step5_refined_similarities.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step5_refined_similarities.pdf Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step6_final_pair_scores.png Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step6_final_pair_scores.pdf Generating validation report... Validation report saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/validation_report_example_16.txt đŸŽ¯ Skipping 4-STAGE example visualizations (--skip_four_stage_viz is set)