shaoncsecu's picture
Experiment-1: Models
cf4cc7a verified
Raw History Blame Contribute Delete
77.9 kB
wandb: Detected [huggingface_hub.inference, openai] in use.
wandb: Use W&B Weave for improved LLM call tracing. Install Weave with `pip install weave` then add `import weave` to the top of your script.
wandb: For more information, check out the docs at: https://weave-docs.wandb.ai/
Initial GPU Memory Stats:
Allocated: 0.00 MB
Cached: 0.00 MB
Max Allocated: 0.00 MB
Initializing bidirectional cross-attention model abhinand/MedEmbed-large-v0.1...
============================================================
πŸ¦₯ UNSLOTH MODE: Loading encoder with optimized kernels
============================================================
πŸ¦₯ Unsloth Status:
Available: βœ… Yes
FastSentenceTransformer: βœ… Yes (preferred)
Version: 2026.1.4
CUDA: βœ… Yes
GPU: NVIDIA GeForce RTX 4080 SUPER
GPU Memory: 16.71 GB
API: FastSentenceTransformer (specialized for embeddings) βœ…
Target modules: ['q_proj', 'k_proj', 'v_proj', 'o_proj', 'gate_proj', 'up_proj', 'down_proj']
[INFO] Auto-detected max_seq_length for abhinand/MedEmbed-large-v0.1: 512
============================================================
πŸ¦₯ UNSLOTH MODE: Creating optimized sentence encoder
============================================================
API: FastSentenceTransformer (specialized for embeddings)
Model: abhinand/MedEmbed-large-v0.1
Max seq length: 512
Full finetuning: False
Unsloth: Using fast encoder path for bert (torch.compile + SDPA)
`torch_dtype` is deprecated! Use `dtype` instead!
Unsloth: Enabled gradient checkpointing
⚠️ Unsloth loading failed: Target modules {'down_proj', 'v_proj', 'up_proj', 'k_proj', 'o_proj', 'q_proj', 'gate_proj'} not found in the base model. Please check the target modules and try again.
Traceback (most recent call last):
File "/home/dtim/Ataur/LOKI/Cross_Attention/run_cross_attention.py", line 847, in main
sentence_encoder = create_unsloth_sentence_encoder(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/dtim/Ataur/LOKI/Cross_Attention/unsloth_encoder.py", line 462, in create_unsloth_sentence_encoder
sentence_encoder = FastSentenceTransformer.get_peft_model(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/unsloth/models/sentence_transformer.py", line 1685, in get_peft_model
peft_model = peft_get_peft_model(inner_model, lora_config)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/mapping_func.py", line 122, in get_peft_model
return MODEL_TYPE_TO_PEFT_MODEL_MAPPING[peft_config.task_type](
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/peft_model.py", line 2966, in __init__
super().__init__(model, peft_config, adapter_name, **kwargs)
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/peft_model.py", line 129, in __init__
self.base_model = cls(model, {adapter_name: peft_config}, adapter_name)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/tuners/tuners_utils.py", line 298, in __init__
self.inject_adapter(self.model, adapter_name, low_cpu_mem_usage=low_cpu_mem_usage, state_dict=state_dict)
File "/home/dtim/anaconda3/envs/LOKI/lib/python3.12/site-packages/peft/tuners/tuners_utils.py", line 880, in inject_adapter
raise ValueError(error_msg)
ValueError: Target modules {'down_proj', 'v_proj', 'up_proj', 'k_proj', 'o_proj', 'q_proj', 'gate_proj'} not found in the base model. Please check the target modules and try again.
Falling back to standard SentenceTransformer loading...
Encoder fine-tuning enabled (trainable_encoder=True). No encoder LoRA adapters are attached by default.
Initialized model without Flash Attention
Embedding dimension: 1024
Max sequence length: 512
Loading datasets...
Loaded 19516 training examples
⚑ Sampled 10000 training examples (deterministic, seed=42)
Loaded 2444 evaluation examples
⚑ Sampled 1000 evaluation examples (deterministic)
Using top_k_sparse attention mechanism
Initializing Top-K Sparse Attention (k=5)
Initializing top-k sparse attention with method: zeros
Successfully applied zeros initialization to top-k sparse attention
Initializing top-k sparse attention with method: zeros
Successfully applied zeros initialization to top-k sparse attention
Using cosine pair scoring method
Bidirectional model initialized with top_k=5, pair_score_method=cosine, share_weights=True, use_refinement=False
Using initialization method: zeros
Initialization parameters: {'bias_value': 0.0}
Sentence encoder is trainable
Sentence encoder dtype detected: torch.float32
Model initialized with standard cross-attention layers
πŸ“Š Model Parameter Statistics:
Total parameters: 354,040,343
Trainable parameters: 354,040,343
Frozen parameters: 0
Trainable percentage: 100.00%
🎯 Loss and Aggregation Configuration:
Architecture: Bidirectional
Loss type: bidirectional_triplet
Aggregation method: top_k_pairs
Top-k value: 5
Norm type: rmsnorm
Q/K RMSNorm: False
Embedding caching: ❌ Disabled
Triplet batching: πŸ”’ Isolated examples
Triplet strategy: LIMITED
Max triplets/example: 2
🎲 Initialization Configuration:
Method: zeros
Description: Initialize all weights to zero (tests uniform attention baseline)
Parameters: {'bias_value': 0.0}
πŸ“Š Loss Component Weights (will be normalized):
Triplet weight: 0.5
Attention weight: 0.0
Pair weight: 0.3
βš™οΈ Loss Parameters:
Margin: 0.3
Scale: 10.0
Pair margin: 0.3
Pair score method: cosine
Share attention weights: True
Join path extraction: True
Join path threshold: 0.15
Self-attention enabled: False
Attention mechanism: top_k_sparse
Sparse top-k: 5
🎯 Attention Mechanism Configuration:
Attention type: top_k_sparse
Sparse top-k: 5
Initialization method: zeros
πŸ”„ Bidirectional-Specific Configuration:
Pair margin: 0.3
Pair score method: cosine
Share attention weights: True
Join path extraction: True
Join path threshold: 0.15
Self-attention enabled: False
Output will be saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945
Training configuration saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_config.json
Starting Model Training...
==================================================
Starting training with ID-based triplet batches...
🎯 Training curves tracker initialized:
πŸ“ Output directory: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945
πŸ“Š Tracking batch losses: True
πŸ“ˆ Tracking validation loss: True
πŸ” Tracking row-sentence metrics: True
πŸ’Ύ Auto-save: True
🎨 Auto-plot: False
🎯 Training curves tracking enabled
πŸ† Test-based model saving enabled - will save best models by test metrics
πŸ” Loading row-sentence evaluation data...
πŸ“Š Row-sentence evaluation data:
Test examples: 36
Annotations: 0
Examples with annotations: 0
ℹ️ Row-sentence data in MIMIC format - will sync with MIMIC loader
πŸ” Loading MIMIC evaluation data...
[INFO] Loaded 18 MIMIC annotations
Diagnosis row groundings: 114
Medication row groundings: 162
[INFO] MIMIC evaluation data loaded:
Test examples: 36
Annotations: 18 admissions
Matched examples: 36
βœ… MIMIC evaluation enabled (all 36 test examples)
πŸ”„ Cache disabled: MIMIC test will be re-encoded each epoch
πŸ”— Syncing generic row-sentence evaluation with loaded MIMIC data for consistent AP tracking...
Preparing triplet batches...
Triplet sampling strategy: LIMITED
Max triplets per example: 2
Using ISOLATED triplet batching (triplets from the same example stay together)
Processing examples (isolated batches): 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [00:01<00:00, 5335.99it/s]
πŸ“Š Triplet Generation Statistics (Isolated Batching):
Total examples: 10000
Total triplets: 16502
Avg triplets/example: 1.65
Min triplets/example: 1
Max triplets/example: 2
Examples with 0 triplets: 0
Created 10000 triplet batches for training (all triplets included, adaptive batching)
Scheduler configured for 200000 steps with 20000 warmup steps
πŸ”₯ STAGE 0: FROZEN ENCODER ONLY EVALUATION...
Evaluating with ONLY the frozen sentence encoder (no cross-attention at all)...
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:05<00:00, 4.21s/it]
πŸ”₯ FROZEN ENCODER ONLY Accuracy: 0.806
(Total comparisons: 24296)
πŸ” Stage 0: Evaluating row-sentence alignment...
ℹ️ Using MIMIC evaluator (Level A)...
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.185
Medication: 18 examples, F1=0.242
πŸ”₯ Stage 0 Row-Sent Avg Precision: 0.419
πŸ”₯ Stage 0 Row-Sent Overall Accuracy: 0.213
πŸš€ STAGE 1: SOPHISTICATED MODEL EVALUATION (PRE-TRAINING)...
Evaluating sophisticated model BEFORE training...
πŸš€ Using CUDA device for evaluation
🎯 TRAINING CONFIGURATION: Starting from Stage 0
πŸ“‹ Configuration detected:
Encoder trainable: True
Use cache: False
Use LoRA: False
Encoder-only training: True
Mode: Encoder-only training (cross-attention heads frozen)
Encoder tuning mode: gradual unfreezing
πŸ”“ Gradual unfreezing initialized: 2/24 top layers trainable (layers 22..23)
Enabling gradient checkpointing for memory efficiency
Optimizer groups configured: head=0 @ 1.00e-04, encoder=391 @ 1.00e-05
Initialized EncoderOnlyTripletLoss with margin=0.3, scale=10.0, ranking=infonce
Initialized EncoderOnlyTripletLoss with margin=0.3, scale=10.0
Using cosine LR schedule with warmup
πŸ”§ Using bidirectional evaluation for bidirectional model
Evaluating bidirectional model with join path extraction...
Evaluating examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:13<00:00, 4.21s/it]
πŸš€ SOPHISTICATED MODEL Accuracy (Pre-training): 0.806
πŸ” Stage 1: Evaluating row-sentence alignment (Sophisticated Untrained)...
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.185
Medication: 18 examples, F1=0.242
πŸš€ Stage 1 Row-Sent Avg Precision: 0.418
πŸš€ Stage 1 Row-Sent Overall Accuracy: 0.213
πŸ“Š ARCHITECTURE BENEFIT:
πŸ”₯β†’πŸš€ Sophisticated model benefit: +0.000 (0.0%)
πŸ“ˆ ROW-SENTENCE PROGRESSION (Avg Precision):
πŸ”₯ Stage 0 (Frozen): 0.419
πŸš€ Stage 1 (Sophisticated): 0.418 (+-0.000)
⚠️ WARNING: Sophisticated model barely improves over frozen encoder!
Consider if the architecture complexity is necessary for this task.
πŸ“Š Logging initial stage metrics to wandb...
βœ… Baseline metrics logged to wandb (epoch 0)
πŸ“Š Adding initial stage metrics to training curves...
πŸ“Š Initial stage metrics recorded:
πŸ”₯ Stage 0 (Frozen): 0.81 acc, 0.42 row-sent AP
πŸš€ Stage 1 (Sophisticated Untrained): 0.81 acc, 0.42 row-sent AP
🎯 Stage 2 (Initial): 0.00 acc, 0.00 row-sent AP
πŸ“Š Adding Epoch 0 (untrained model) baseline to training curves...
πŸŽ“ STARTING TRAINING OF ENCODER ONLY (NO HEADS)
Training model type: Encoder Only (No Heads)
Training from Stage: 0
==================== Epoch 1/20 ====================
Training model has 26242048 trainable parameters
Epoch 1: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:48:51<00:00, 1.73s/it, batch_loss=0.687]
Current learning rates: 5.00e-06 | margin=0.300
Epoch 1 loss statistics - Mean: 0.689, Min: 0.520, Max: 0.854
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:57<00:00, 4.20s/it]
Epoch 1 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36305.59it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.58it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.203
Medication: 18 examples, F1=0.244
Row-Sent Overall Acc: 0.224
Row-Sent Avg Precision: 0.435
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37328.79it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.44it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.203
Medication: 18 examples, F1=0.244
MIMIC Diagnosis - P: 0.119, R: 0.847, F1: 0.203
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244
πŸ“Š Epoch 1 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21544.8s
πŸ“ˆ Learning Rate: 5.00e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.44
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
New best model! Accuracy: 0.806
New best test overall accuracy! 0.224 (Epoch 1)
New best test average precision! 0.435 (Epoch 1)
==================== Epoch 2/20 ====================
πŸ”“ Gradual unfreezing: enabled layers 21..21 (total 3/24)
Epoch 2: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:49:35<00:00, 1.74s/it, batch_loss=0.683]
Current learning rates: 1.00e-05 | margin=0.300
Epoch 2 loss statistics - Mean: 0.689, Min: 0.523, Max: 0.871
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:21<00:00, 4.22s/it]
Epoch 2 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36314.32it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.60it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.203
Medication: 18 examples, F1=0.244
Row-Sent Overall Acc: 0.223
Row-Sent Avg Precision: 0.435
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37748.74it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.77it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.203
Medication: 18 examples, F1=0.244
MIMIC Diagnosis - P: 0.118, R: 0.847, F1: 0.203
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244
πŸ“Š Epoch 2 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21613.4s
πŸ“ˆ Learning Rate: 1.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.44
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 3/20 ====================
πŸ”“ Gradual unfreezing: enabled layers 20..20 (total 4/24)
Epoch 3: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:32<00:00, 1.74s/it, batch_loss=0.683]
Current learning rates: 9.92e-06 | margin=0.300
Epoch 3 loss statistics - Mean: 0.689, Min: 0.504, Max: 0.875
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:17<00:00, 4.22s/it]
Epoch 3 Accuracy: 0.805
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 35511.51it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.69it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.435
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36846.01it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.21it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.244
MIMIC Diagnosis - P: 0.117, R: 0.847, F1: 0.201
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244
πŸ“Š Epoch 3 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21666.8s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.44
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 4/20 ====================
πŸ”“ Gradual unfreezing: enabled layers 19..19 (total 5/24)
Epoch 4: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:41<00:00, 1.74s/it, batch_loss=0.683]
Current learning rates: 9.70e-06 | margin=0.300
Epoch 4 loss statistics - Mean: 0.689, Min: 0.525, Max: 0.854
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:11<00:00, 4.21s/it]
Epoch 4 Accuracy: 0.805
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36756.32it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.76it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.244
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.435
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 18810.88it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 56.02it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.244
MIMIC Diagnosis - P: 0.117, R: 0.847, F1: 0.201
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244
πŸ“Š Epoch 4 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21669.1s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 5/20 ====================
πŸ”“ Gradual unfreezing: enabled layers 18..18 (total 6/24)
Epoch 5: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:47<00:00, 1.74s/it, batch_loss=0.686]
Current learning rates: 9.33e-06 | margin=0.300
Epoch 5 loss statistics - Mean: 0.689, Min: 0.535, Max: 0.854
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:56<00:00, 4.20s/it]
Epoch 5 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36054.19it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.85it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.434
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 39097.60it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.19it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
MIMIC Diagnosis - P: 0.117, R: 0.849, F1: 0.201
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.243
πŸ“Š Epoch 5 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21660.5s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 6/20 ====================
Epoch 6: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:51<00:00, 1.75s/it, batch_loss=0.688]
Current learning rates: 8.83e-06 | margin=0.300
Epoch 6 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.857
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:18<00:00, 4.22s/it]
Epoch 6 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 35603.62it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.66it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.244
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.434
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36756.32it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 58.97it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.244
MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.201
MIMIC Medication - P: 0.153, R: 0.840, F1: 0.244
πŸ“Š Epoch 6 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21686.2s
πŸ“ˆ Learning Rate: 8.83e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 7/20 ====================
Epoch 7: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:28<00:00, 1.75s/it, batch_loss=0.682]
Current learning rates: 8.21e-06 | margin=0.300
Epoch 7 loss statistics - Mean: 0.689, Min: 0.512, Max: 0.866
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:17<00:00, 4.22s/it]
Epoch 7 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37504.95it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.92it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.16it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.435
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 16710.37it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 57.66it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.201
MIMIC Medication - P: 0.152, R: 0.840, F1: 0.243
πŸ“Š Epoch 7 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21722.1s
πŸ“ˆ Learning Rate: 8.21e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 8/20 ====================
Epoch 8: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:15<00:00, 1.75s/it, batch_loss=0.682]
Current learning rates: 7.50e-06 | margin=0.300
Epoch 8 loss statistics - Mean: 0.689, Min: 0.536, Max: 0.865
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:14<00:00, 4.21s/it]
Epoch 8 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37099.49it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.28it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.201
Medication: 18 examples, F1=0.243
Row-Sent Overall Acc: 0.222
Row-Sent Avg Precision: 0.434
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36587.10it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.93it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.243
MIMIC Diagnosis - P: 0.117, R: 0.851, F1: 0.200
MIMIC Medication - P: 0.152, R: 0.840, F1: 0.243
πŸ“Š Epoch 8 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21706.6s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 9/20 ====================
Epoch 9: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:10<00:00, 1.75s/it, batch_loss=0.681]
Current learning rates: 6.71e-06 | margin=0.300
Epoch 9 loss statistics - Mean: 0.689, Min: 0.522, Max: 0.859
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:14<00:00, 4.21s/it]
Epoch 9 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 33149.27it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.62it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
Row-Sent Overall Acc: 0.221
Row-Sent Avg Precision: 0.434
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 35704.65it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 58.34it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200
MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241
πŸ“Š Epoch 9 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 1)
⏱️ Epoch Time: 21701.2s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 10/20 ====================
Epoch 10: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:09<00:00, 1.75s/it, batch_loss=0.678]
Current learning rates: 5.87e-06 | margin=0.300
Epoch 10 loss statistics - Mean: 0.689, Min: 0.531, Max: 0.860
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:17<00:00, 4.22s/it]
Epoch 10 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 34735.44it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.24it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
Row-Sent Overall Acc: 0.221
Row-Sent Avg Precision: 0.434
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 38101.17it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.52it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200
MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241
πŸ“Š Epoch 10 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21704.0s
πŸ“ˆ Learning Rate: 5.87e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
New best model! Accuracy: 0.806
==================== Epoch 11/20 ====================
Epoch 11: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:28<00:00, 1.75s/it, batch_loss=0.678]
Current learning rates: 5.00e-06 | margin=0.300
Epoch 11 loss statistics - Mean: 0.689, Min: 0.530, Max: 0.859
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:21<00:00, 4.22s/it]
Epoch 11 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 35229.80it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.53it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
Row-Sent Overall Acc: 0.221
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 34976.82it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.47it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.17it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.241
MIMIC Diagnosis - P: 0.116, R: 0.851, F1: 0.200
MIMIC Medication - P: 0.151, R: 0.840, F1: 0.241
πŸ“Š Epoch 11 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21726.3s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 12/20 ====================
Epoch 12: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:23<00:00, 1.75s/it, batch_loss=0.685]
Current learning rates: 4.13e-06 | margin=0.300
Epoch 12 loss statistics - Mean: 0.689, Min: 0.508, Max: 0.865
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:16<00:00, 4.22s/it]
Epoch 12 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 33391.19it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.82it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.220
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37588.98it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.78it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.11it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.200
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.116, R: 0.847, F1: 0.200
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 12 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21717.2s
πŸ“ˆ Learning Rate: 4.13e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 13/20 ====================
Epoch 13: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:05<00:00, 1.75s/it, batch_loss=0.684]
Current learning rates: 3.29e-06 | margin=0.300
Epoch 13 loss statistics - Mean: 0.689, Min: 0.528, Max: 0.890
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:57<00:00, 4.20s/it]
Epoch 13 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36711.63it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.63it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 39808.84it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.23it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 13 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21679.7s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 14/20 ====================
Epoch 14: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:45<00:00, 1.74s/it, batch_loss=0.686]
Current learning rates: 2.50e-06 | margin=0.300
Epoch 14 loss statistics - Mean: 0.689, Min: 0.542, Max: 0.859
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:57<00:00, 4.20s/it]
Epoch 14 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36114.55it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.32it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 38796.23it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.26it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 14 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21660.1s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 15/20 ====================
Epoch 15: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:47<00:00, 1.74s/it, batch_loss=0.688]
Current learning rates: 1.79e-06 | margin=0.300
Epoch 15 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.869
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:03<00:00, 4.20s/it]
Epoch 15 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37356.49it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.65it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36428.21it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.69it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 15 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21666.9s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 16/20 ====================
Epoch 16: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:44<00:00, 1.74s/it, batch_loss=0.689]
Current learning rates: 1.17e-06 | margin=0.300
Epoch 16 loss statistics - Mean: 0.689, Min: 0.534, Max: 0.860
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:57<00:00, 4.20s/it]
Epoch 16 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 36393.09it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.83it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 39250.05it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.32it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 16 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21658.9s
πŸ“ˆ Learning Rate: 1.17e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 17/20 ====================
Epoch 17: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:45<00:00, 1.74s/it, batch_loss=0.686]
Current learning rates: 1.00e-06 | margin=0.300
Epoch 17 loss statistics - Mean: 0.689, Min: 0.534, Max: 0.849
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:57<00:00, 4.20s/it]
Epoch 17 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37467.73it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.80it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37035.80it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.21it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 17 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21660.2s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 18/20 ====================
Epoch 18: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:43<00:00, 1.74s/it, batch_loss=0.679]
Current learning rates: 1.00e-06 | margin=0.300
Epoch 18 loss statistics - Mean: 0.689, Min: 0.532, Max: 0.872
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:09:58<00:00, 4.20s/it]
Epoch 18 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 37090.38it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.83it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 34647.76it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.39it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.19it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 18 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21658.0s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 19/20 ====================
Epoch 19: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:50:59<00:00, 1.75s/it, batch_loss=0.687]
Current learning rates: 1.00e-06 | margin=0.300
Epoch 19 loss statistics - Mean: 0.689, Min: 0.518, Max: 0.866
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:17<00:00, 4.22s/it]
Epoch 19 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 32640.50it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.64it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 34799.48it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 52.03it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 19 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21693.0s
πŸ“ˆ Learning Rate: 1.00e-06
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
==================== Epoch 20/20 ====================
Epoch 20: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10000/10000 [4:51:12<00:00, 1.75s/it, batch_loss=0.684]
Current learning rates: 1.00e-06 | margin=0.300
Epoch 20 loss statistics - Mean: 0.689, Min: 0.538, Max: 0.871
Loss not decreasing. Patience: 1/1
Reset learning rate to 5.00e-05
Evaluating after epoch...
πŸ”Ž Validation evaluator: Stage 0 (frozen encoder baseline)
πŸ”₯ Evaluating with FROZEN ENCODER ONLY (no cross-attention, aggregation=top_k_pairs)
Evaluating frozen encoder baseline: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 1000/1000 [1:10:14<00:00, 4.21s/it]
Epoch 20 Accuracy: 0.806
Collecting tables and contexts for test_row_sent_epoch split...
Processing test_row_sent_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 34560.53it/s]
Found 36 unique tables and 39 unique contexts in test_row_sent_epoch split
Using device: cuda for embedding cache
Encoding test_row_sent_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 61.65it/s]
Encoding test_row_sent_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.18it/s]
πŸ”Ž Row-sent evaluator: Stage 0 (frozen encoder baseline)
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
Row-Sent Overall Acc: 0.219
Row-Sent Avg Precision: 0.433
Examples evaluated: 36
Collecting tables and contexts for mimic_test_epoch split...
Processing mimic_test_epoch examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 38333.32it/s]
Found 36 unique tables and 39 unique contexts in mimic_test_epoch split
Using device: cuda for embedding cache
Encoding mimic_test_epoch tables: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:00<00:00, 59.05it/s]
Encoding mimic_test_epoch contexts: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 39/39 [00:07<00:00, 5.14it/s]
[INFO] MIMIC Row Grounding evaluation: 36 examples
Diagnosis: 18 examples, F1=0.198
Medication: 18 examples, F1=0.240
MIMIC Diagnosis - P: 0.115, R: 0.847, F1: 0.198
MIMIC Medication - P: 0.149, R: 0.840, F1: 0.240
πŸ“Š Epoch 20 Summary:
πŸ”₯ Train Loss: 0.69 Β± 0.04
🎯 Val Accuracy: 0.81
πŸ† Best Accuracy: 0.81 (Epoch 10)
⏱️ Epoch Time: 21702.7s
πŸ“ˆ Learning Rate: 5.00e-05
πŸ” Row-Sent Overall Acc: 0.22
πŸ† Best Test Overall Acc: 0.22 (Epoch 1)
πŸ” Row-Sent Avg Precision: 0.43
πŸ† Best Test Avg Precision: 0.44 (Epoch 1)
Stage 0 note: encoder fine-tuning active; eval uses on-the-fly embeddings (no cache)
Loading best model from epoch 10 with accuracy 0.806
Best model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_model_epoch_10/model.pt
🎯 SAVING BEST TEST-BASED MODELS:
Loading best test overall accuracy model from epoch 1 (0.224)
Best test overall accuracy model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_test_overall_acc_epoch_1/model.pt
Loading best test average precision model from epoch 1 (0.435)
Best test average precision model saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/abhinand/MedEmbed-large-v0.1/best_test_avg_precision_epoch_1/model.pt
Reloaded validation-based best model for final analysis
==========================================================================================
🎯 COMPREHENSIVE 3-STAGE IMPACT ANALYSIS
==========================================================================================
πŸ“Š COMPLETE PERFORMANCE BREAKDOWN:
πŸ”₯ Stage 0 - Frozen Encoder Only: 0.806
πŸš€ Stage 1 - Sophisticated (Pre): 0.806 (+0.000)
πŸ† Stage 2 - Trained Model: 0.806 (+0.000)
πŸ“ˆ Total Improvement: +0.000 (0.0%)
πŸ† TEST METRICS SUMMARY:
🎯 Best Test Overall Accuracy: 0.224 (Epoch 1)
🎯 Best Test Average Precision: 0.435 (Epoch 1)
πŸ“Š Validation vs Test Performance:
- Validation Best: 0.806 (Epoch 10)
- Test Overall Acc: 0.224 (Epoch 1)
- Test Avg Precision: 0.435 (Epoch 1)
⚠️ NOTE: Best test overall accuracy occurred at epoch 1, not at best validation epoch 10βœ… Final summary logged to wandb
🎨 Skipping 3-STAGE example visualizations (skip_four_stage_viz=True)
==========================================================================================
======================================================================
🎯 TRAINING CURVES SUMMARY
======================================================================
======================================================================
🎯 TRAINING SUMMARY - abhinand/MedEmbed-large-v0.1
======================================================================
πŸ“Š Total Epochs: 20
πŸ† Best Accuracy: 0.81 (Epoch 10)
πŸ“ˆ Final Accuracy: 0.81
πŸ“‰ Accuracy Improvement: -0.00
πŸ”₯ Initial Loss: 0.69
🎯 Final Loss: 0.69
πŸ“‰ Loss Reduction: -0.00
⏱️ Average Epoch Time: 21674.9s
πŸ• Total Training Time: 433497.5s (7225.0 min)
πŸ† Best Test Metrics:
πŸ” Best Overall Accuracy: 0.22 (Epoch 1)
πŸ” Best Average Precision: 0.44 (Epoch 1)
πŸ“Š Initial Stage Metrics:
πŸ”₯ Stage 0 (Frozen): 0.81 acc, 0.42 row-sent AP
πŸš€ Stage 1 (Sophisticated Untrained): 0.81 acc, 0.42 row-sent AP
🎯 Stage 2 (Initial): 0.00 acc, 0.00 row-sent AP
======================================================================
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_final_training_curves.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_final_training_curves.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_training_loss.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_training_loss.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_validation_accuracy.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_validation_accuracy.pdf
πŸ“ˆ Generating batch-level analysis...
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.pdf
🌑️ Batch losses heatmap saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_heatmap.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.pdf
πŸ“ˆ Batch losses plot for epoch 1 saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_1.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.pdf
πŸ“ˆ Batch losses plot for epoch 21 saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots/abhinand_MedEmbed-large-v0.1_batch_losses_epoch_21.png
βœ… Training curves analysis complete!
πŸ“ All plots saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_plots
πŸ’Ύ Training data saved to: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/training_data
πŸŽ“ TRAINING COMPLETED: Returning trained Encoder Only (No Heads)
βœ… Training completed! The returned model is the BEST performing model from training.
πŸ“ Check the training logs above to see which epoch achieved the highest validation accuracy.
Saved best model to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt
Evaluating on Test Set...
Loaded 36 test examples
Cache disabled - evaluation will compute embeddings on-the-fly
Evaluating bidirectional model with join path extraction...
Evaluating examples: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 36/36 [00:18<00:00, 1.94it/s]
Saved 36 join path examples to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/extracted_join_paths.json
Test Accuracy: 0.931
Join Paths Extracted: 36
Evaluation results saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/test_metrics.json
🎨 Generating visualizations using the BEST model (highest validation accuracy)...
βœ… CONFIRMED: Using BEST model checkpoint from training
πŸ“ Best model path: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt
πŸ”§ Using aggregation method: top_k_pairs
βœ… Best model file confirmed to exist
Running comprehensive visualizations using top_k_pairs...
πŸ” Loading trained model from: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt
πŸ“ Loading TRAINED model from checkpoint: output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt
βœ… CONFIRMED: Loading BEST model checkpoint (highest validation accuracy)
πŸ“„ Found configuration file: args.json
βœ… Loaded from config: model_type=Bidirectional, embedding_dim=1024, attention_type=top_k_sparse
Creating dynamic encoder for 1024 dimensions on cuda
πŸ“ Final embedding dimension: 1024
Creating BidirectionalTableTextModel...
πŸš€ Using top_k_sparse attention mechanism
πŸ“‹ Using top-5 sparse attention
Using top_k_sparse attention mechanism
Initializing Top-K Sparse Attention (k=5)
Initializing top-k sparse attention with method: xavier_uniform
Successfully applied xavier_uniform initialization to top-k sparse attention
Initializing top-k sparse attention with method: xavier_uniform
Successfully applied xavier_uniform initialization to top-k sparse attention
Using cosine pair scoring method
Bidirectional model initialized with top_k=3, pair_score_method=cosine, share_weights=False, use_refinement=True
Using initialization method: xavier_uniform
Sentence encoder is frozen
Sentence encoder dtype detected: torch.float32
Loading trained weights from output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/MedEmbed-large-v0.1_best.pt
βœ… Trained weights loaded successfully (custom layers only)
πŸ† VERIFICATION: Best model weights are now active for visualization
🎯 Model ready for visualization: BEST TRAINED model on cuda
βœ… Trained model loaded with dynamic dimension detection on cuda
Processing example 16 (ID: 2221973625275923611):
8 rows, 69 sentences
Using aggregation method: top_k_pairs
Generating comprehensive 4-panel analysis...
πŸ” Creating comprehensive analysis for Example 16...
Found 8 rows and 69 sentences
Computing comprehensive similarities...
Computing raw embedding similarities...
Computing bidirectional attention and contextualized similarities...
Applying top_k_sparse attention mechanism...
Computing contextualized similarities...
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.pdf
πŸ’Ύ Saved comprehensive analysis to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png
πŸ“Š Analysis Summary for Example 16:
Raw Embeddings - Range: [-0.101, 0.082]
Contextualized Similarities - Range: [-0.071, 0.396]
Final Model Similarities - Range: [-0.091, 0.082]
Cross-Attention - Range: [0.000, 0.200]
πŸ” Top 3 pairs by Raw Embeddings:
1. Row 2 - Sentence 57: 0.082
2. Row 6 - Sentence 55: 0.079
3. Row 6 - Sentence 45: 0.074
πŸ” Top 3 pairs by Contextualized Similarities:
1. Row 1 - Sentence 3: 0.396
2. Row 3 - Sentence 2: 0.379
3. Row 5 - Sentence 4: 0.377
πŸ” Top 3 pairs by Final Model Similarities:
1. Row 4 - Sentence 27: 0.082
2. Row 5 - Sentence 20: 0.082
3. Row 4 - Sentence 10: 0.079
πŸ” Top 3 pairs by Cross-Attention:
1. Row 5 - Sentence 2: 0.200
2. Row 5 - Sentence 3: 0.200
3. Row 5 - Sentence 4: 0.200
πŸ“„ Detailed analysis saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/example_16_similarity_analysis.txt
All visualizations saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations
πŸ”¬ Creating comprehensive model analysis...
πŸ”¬ Creating comprehensive analysis for example 16 using BEST model...
πŸ” Creating comprehensive analysis for Example 16...
Found 8 rows and 69 sentences
Computing comprehensive similarities...
Computing raw embedding similarities...
Computing bidirectional attention and contextualized similarities...
Applying top_k_sparse attention mechanism...
Computing contextualized similarities...
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.pdf
πŸ’Ύ Saved comprehensive analysis to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/comprehensive_analysis_example_16.png
πŸ“Š Analysis Summary for Example 16:
Raw Embeddings - Range: [0.456, 0.725]
Contextualized Similarities - Range: [0.834, 0.949]
Final Model Similarities - Range: [0.456, 0.725]
Cross-Attention - Range: [0.000, 0.200]
πŸ” Top 3 pairs by Raw Embeddings:
1. Row 8 - Sentence 25: 0.725
2. Row 6 - Sentence 25: 0.723
3. Row 2 - Sentence 28: 0.717
πŸ” Top 3 pairs by Contextualized Similarities:
1. Row 4 - Sentence 38: 0.949
2. Row 1 - Sentence 2: 0.943
3. Row 4 - Sentence 4: 0.943
πŸ” Top 3 pairs by Final Model Similarities:
1. Row 8 - Sentence 25: 0.725
2. Row 6 - Sentence 25: 0.723
3. Row 2 - Sentence 28: 0.717
πŸ” Top 3 pairs by Cross-Attention:
1. Row 5 - Sentence 2: 0.200
2. Row 5 - Sentence 3: 0.200
3. Row 5 - Sentence 4: 0.200
πŸ“„ Detailed analysis saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/example_16_similarity_analysis.txt
πŸ”¬ Generating step-by-step diagnostics for example 16...
πŸ”¬ Generating step-by-step diagnostics for example 16...
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step1_raw_similarities.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step1_raw_similarities.pdf
Applying top_k_sparse attention mechanism...
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step2_forward_attention.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step2_forward_attention.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step3_reverse_attention.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step3_reverse_attention.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step4_contextualized_similarities.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step4_contextualized_similarities.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step5_refined_similarities.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step5_refined_similarities.pdf
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step6_final_pair_scores.png
Saved plot to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/step6_final_pair_scores.pdf
Generating validation report...
Validation report saved to output_best_model/bidirectional_MedEmbed-large-v0.1_20260128_011945/visualizations/diagnostics_example_16/validation_report_example_16.txt
🎯 Skipping 4-STAGE example visualizations (--skip_four_stage_viz is set)