test / artifacts /logs /trainer.log
XProger's picture
trainer-runtime e2e job_e2e_1774222987861 (metadata)
ac3cfe2 verified
Raw
History Blame Contribute Delete
9.78 kB
2026-03-22 23:43:09,773 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:09,774 [INFO] ==> config loaded
2026-03-22 23:43:09,776 [INFO] ==> job_name: trainer-e2e-job_e2e_1774222987861
2026-03-22 23:43:09,777 [INFO] ==> job_id: job_e2e_1774222987861
2026-03-22 23:43:09,777 [INFO] ==> config source: http://host.docker.internal:18787/api/v1/trainer/jobs/job_e2e_1774222987861/bootstrap?token=cfg_836915adce2f337d2264246143c15c255d9ae74b10e6f126
2026-03-22 23:43:09,800 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:10,074 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
2026-03-22 23:43:10,249 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
2026-03-22 23:43:10,250 [INFO] ==> Hugging Face upload ready for account: XProger
2026-03-22 23:43:10,252 [INFO] ==> preparing assets
2026-03-22 23:43:10,252 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:19,148 [INFO] ==> starting training
2026-03-22 23:43:19,150 [INFO] ==> loading model: /app
2026-03-22 23:43:19,151 [INFO] ==> logical base model id: Qwen/Qwen2.5-7B-Instruct
2026-03-22 23:43:19,151 [INFO] ==> method: qlora
2026-03-22 23:43:19,152 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:19,459 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
2026-03-22 23:43:19,942 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
2026-03-22 23:43:28,508 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:29,000 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/json/json.py "HTTP/1.1 200 OK"
2026-03-22 23:43:32,143 [INFO] Unsloth: Padding-free batching auto-enabled for SFTTrainer instance.
2026-03-22 23:43:32,334 [WARNING] num_proc must be <= 2. Reducing num_proc to 2 for dataset of size 2.
2026-03-22 23:43:37,654 [WARNING] num_proc must be <= 1. Reducing num_proc to 1 for dataset of size 1.
2026-03-22 23:43:41,594 [INFO] ==> training started
2026-03-22 23:43:42,839 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:50,075 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:50,729 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:50,963 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:51,219 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:51,348 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:51,568 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:51,580 [INFO] ==> training finished
2026-03-22 23:43:51,580 [INFO] ==> saving lora adapters to /output/job_e2e_1774222987861/lora/trainer-e2e-job_e2e_1774222987861
2026-03-22 23:43:51,595 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:43:51,612 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:51,630 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:51,766 [INFO] ==> saving merged model to /output/job_e2e_1774222987861/merged/trainer-e2e-job_e2e_1774222987861
2026-03-22 23:43:51,766 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:43:51,766 [INFO] ==> merged 16-bit save may take several minutes for 7B model
2026-03-22 23:44:16,075 [INFO] ==> merged model saved
2026-03-22 23:44:16,079 [INFO] ==> training artifacts saved
2026-03-22 23:44:16,079 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:16,079 [INFO] ==> training finished
2026-03-22 23:44:16,088 [INFO] ==> starting evaluation
2026-03-22 23:44:16,132 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:22,174 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:22,232 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
2026-03-22 23:44:22,356 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:22,357 [INFO] ==> evaluation finished
2026-03-22 23:44:22,403 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:22,469 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar from /output/job_e2e_1774222987861/lora/trainer-e2e-job_e2e_1774222987861 (mode=w, files=6, size_bytes=21538633)
2026-03-22 23:44:22,500 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar
2026-03-22 23:44:22,500 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/lora: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar.gz
2026-03-22 23:44:22,502 [ERROR] upload step failed: lora_archive
Traceback (most recent call last):
File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
result = operation()
^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 155, in <lambda>
lambda: self._archive_and_upload_dir(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 64, in _archive_and_upload_dir
self.archiver.upload_archive(
File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
with open(archive_path, "rb") as f:
^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar.gz'
2026-03-22 23:44:22,504 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar from /output/job_e2e_1774222987861/merged/trainer-e2e-job_e2e_1774222987861 (mode=w, files=9, size_bytes=15242728923)
2026-03-22 23:44:33,444 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar
2026-03-22 23:44:33,444 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/merged: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar.gz
2026-03-22 23:44:33,445 [ERROR] upload step failed: merged_archive
Traceback (most recent call last):
File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
result = operation()
^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 169, in <lambda>
lambda: self._archive_and_upload_dir(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 64, in _archive_and_upload_dir
self.archiver.upload_archive(
File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
with open(archive_path, "rb") as f:
^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar.gz'
2026-03-22 23:44:33,449 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar from /output/job_e2e_1774222987861 (mode=w, files=36, size_bytes=15291110697)
2026-03-22 23:44:44,091 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar
2026-03-22 23:44:44,091 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/full-archive: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar.gz
2026-03-22 23:44:44,091 [ERROR] upload step failed: full_archive
Traceback (most recent call last):
File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
result = operation()
^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 182, in <lambda>
lambda: self._archive_and_upload_dir(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/trainer/app/pipeline/upload_runner.py", line 64, in _archive_and_upload_dir
self.archiver.upload_archive(
File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
with open(archive_path, "rb") as f:
^^^^^^^^^^^^^^^^^^^^^^^^
FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar.gz'
2026-03-22 23:44:44,136 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2026-03-22 23:44:44,391 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
2026-03-22 23:45:21,012 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"
2026-03-22 23:45:21,184 [INFO] HTTP Request: POST https://huggingface.co/XProger/test.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
2026-03-22 23:45:21,438 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/xet-write-token/main "HTTP/1.1 200 OK"
2026-03-22 23:45:48,946 [WARNING] No files have been modified since last commit. Skipping to prevent empty commit.
2026-03-22 23:45:49,203 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/revision/main "HTTP/1.1 200 OK"
2026-03-22 23:45:49,384 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
2026-03-22 23:45:49,567 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"