XProger commited on
Commit
ac3cfe2
·
verified ·
1 Parent(s): 0e1cc3f

trainer-runtime e2e job_e2e_1774222987861 (metadata)

Browse files
Files changed (1) hide show
  1. artifacts/logs/trainer.log +74 -74
artifacts/logs/trainer.log CHANGED
@@ -1,57 +1,57 @@
1
- 2026-03-22 23:33:47,924 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2
- 2026-03-22 23:33:47,926 [INFO] ==> config loaded
3
- 2026-03-22 23:33:47,926 [INFO] ==> job_name: trainer-e2e-job_e2e_1774222425841
4
- 2026-03-22 23:33:47,926 [INFO] ==> job_id: job_e2e_1774222425841
5
- 2026-03-22 23:33:47,927 [INFO] ==> config source: http://host.docker.internal:18787/api/v1/trainer/jobs/job_e2e_1774222425841/bootstrap?token=cfg_6effc79d0f3d1d0f0c434f2e57359528dc7f23ed86b34957
6
- 2026-03-22 23:33:47,951 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
7
- 2026-03-22 23:33:48,233 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
8
- 2026-03-22 23:33:48,409 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
9
- 2026-03-22 23:33:48,409 [INFO] ==> Hugging Face upload ready for account: XProger
10
- 2026-03-22 23:33:48,411 [INFO] ==> preparing assets
11
- 2026-03-22 23:33:48,411 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
12
- 2026-03-22 23:33:57,748 [INFO] ==> starting training
13
- 2026-03-22 23:33:57,753 [INFO] ==> loading model: /app
14
- 2026-03-22 23:33:57,753 [INFO] ==> logical base model id: Qwen/Qwen2.5-7B-Instruct
15
- 2026-03-22 23:33:57,753 [INFO] ==> method: qlora
16
- 2026-03-22 23:33:57,755 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
17
- 2026-03-22 23:33:58,016 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
18
- 2026-03-22 23:33:58,485 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
19
- 2026-03-22 23:34:07,674 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
20
- 2026-03-22 23:34:08,118 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/json/json.py "HTTP/1.1 200 OK"
21
- 2026-03-22 23:34:11,309 [INFO] Unsloth: Padding-free batching auto-enabled for SFTTrainer instance.
22
- 2026-03-22 23:34:11,504 [WARNING] num_proc must be <= 2. Reducing num_proc to 2 for dataset of size 2.
23
- 2026-03-22 23:34:16,779 [WARNING] num_proc must be <= 1. Reducing num_proc to 1 for dataset of size 1.
24
- 2026-03-22 23:34:20,429 [INFO] ==> training started
25
- 2026-03-22 23:34:21,583 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
26
- 2026-03-22 23:34:28,666 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
27
- 2026-03-22 23:34:29,302 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
28
- 2026-03-22 23:34:29,526 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
29
- 2026-03-22 23:34:29,783 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
30
- 2026-03-22 23:34:29,906 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
31
- 2026-03-22 23:34:30,110 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
32
- 2026-03-22 23:34:30,120 [INFO] ==> training finished
33
- 2026-03-22 23:34:30,121 [INFO] ==> saving lora adapters to /output/job_e2e_1774222425841/lora/trainer-e2e-job_e2e_1774222425841
34
- 2026-03-22 23:34:30,134 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
35
- 2026-03-22 23:34:30,156 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
36
- 2026-03-22 23:34:30,175 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
37
- 2026-03-22 23:34:30,293 [INFO] ==> saving merged model to /output/job_e2e_1774222425841/merged/trainer-e2e-job_e2e_1774222425841
38
- 2026-03-22 23:34:30,294 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
39
- 2026-03-22 23:34:30,295 [INFO] ==> merged 16-bit save may take several minutes for 7B model
40
- 2026-03-22 23:34:51,867 [INFO] ==> merged model saved
41
- 2026-03-22 23:34:51,871 [INFO] ==> training artifacts saved
42
- 2026-03-22 23:34:51,871 [INFO] ==> training finished
43
- 2026-03-22 23:34:51,871 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
44
- 2026-03-22 23:34:51,879 [INFO] ==> starting evaluation
45
- 2026-03-22 23:34:52,054 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
46
- 2026-03-22 23:34:58,012 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
47
- 2026-03-22 23:34:58,032 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
48
- 2026-03-22 23:34:58,133 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
49
- 2026-03-22 23:34:58,135 [INFO] ==> evaluation finished
50
- 2026-03-22 23:34:58,171 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
51
- 2026-03-22 23:34:58,239 [INFO] ==> creating archive /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.lora.tar from /output/job_e2e_1774222425841/lora/trainer-e2e-job_e2e_1774222425841 (mode=w, files=6, size_bytes=21538633)
52
- 2026-03-22 23:34:58,260 [INFO] ==> archive ready: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.lora.tar
53
- 2026-03-22 23:34:58,261 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/lora: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.lora.tar.gz
54
- 2026-03-22 23:34:58,261 [ERROR] upload step failed: lora_archive
55
  Traceback (most recent call last):
56
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
57
  result = operation()
@@ -64,11 +64,11 @@ Traceback (most recent call last):
64
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
65
  with open(archive_path, "rb") as f:
66
  ^^^^^^^^^^^^^^^^^^^^^^^^
67
- FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.lora.tar.gz'
68
- 2026-03-22 23:34:58,264 [INFO] ==> creating archive /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.merged.tar from /output/job_e2e_1774222425841/merged/trainer-e2e-job_e2e_1774222425841 (mode=w, files=9, size_bytes=15242728923)
69
- 2026-03-22 23:35:08,940 [INFO] ==> archive ready: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.merged.tar
70
- 2026-03-22 23:35:08,941 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/merged: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.merged.tar.gz
71
- 2026-03-22 23:35:08,942 [ERROR] upload step failed: merged_archive
72
  Traceback (most recent call last):
73
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
74
  result = operation()
@@ -81,11 +81,11 @@ Traceback (most recent call last):
81
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
82
  with open(archive_path, "rb") as f:
83
  ^^^^^^^^^^^^^^^^^^^^^^^^
84
- FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.merged.tar.gz'
85
- 2026-03-22 23:35:08,952 [INFO] ==> creating archive /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.full.tar from /output/job_e2e_1774222425841 (mode=w, files=36, size_bytes=15291110697)
86
- 2026-03-22 23:35:19,191 [INFO] ==> archive ready: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.full.tar
87
- 2026-03-22 23:35:19,192 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/full-archive: /output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.full.tar.gz
88
- 2026-03-22 23:35:19,192 [ERROR] upload step failed: full_archive
89
  Traceback (most recent call last):
90
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
91
  result = operation()
@@ -98,13 +98,13 @@ Traceback (most recent call last):
98
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
99
  with open(archive_path, "rb") as f:
100
  ^^^^^^^^^^^^^^^^^^^^^^^^
101
- FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222425841/trainer-e2e-job_e2e_1774222425841.full.tar.gz'
102
- 2026-03-22 23:35:19,238 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
103
- 2026-03-22 23:35:19,502 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
104
- 2026-03-22 23:35:56,149 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"
105
- 2026-03-22 23:35:56,322 [INFO] HTTP Request: POST https://huggingface.co/XProger/test.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
106
- 2026-03-22 23:35:56,508 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/xet-write-token/main "HTTP/1.1 200 OK"
107
- 2026-03-22 23:36:19,941 [WARNING] No files have been modified since last commit. Skipping to prevent empty commit.
108
- 2026-03-22 23:36:20,218 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/revision/main "HTTP/1.1 200 OK"
109
- 2026-03-22 23:36:20,405 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
110
- 2026-03-22 23:36:20,596 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"
 
1
+ 2026-03-22 23:43:09,773 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
2
+ 2026-03-22 23:43:09,774 [INFO] ==> config loaded
3
+ 2026-03-22 23:43:09,776 [INFO] ==> job_name: trainer-e2e-job_e2e_1774222987861
4
+ 2026-03-22 23:43:09,777 [INFO] ==> job_id: job_e2e_1774222987861
5
+ 2026-03-22 23:43:09,777 [INFO] ==> config source: http://host.docker.internal:18787/api/v1/trainer/jobs/job_e2e_1774222987861/bootstrap?token=cfg_836915adce2f337d2264246143c15c255d9ae74b10e6f126
6
+ 2026-03-22 23:43:09,800 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
7
+ 2026-03-22 23:43:10,074 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
8
+ 2026-03-22 23:43:10,249 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
9
+ 2026-03-22 23:43:10,250 [INFO] ==> Hugging Face upload ready for account: XProger
10
+ 2026-03-22 23:43:10,252 [INFO] ==> preparing assets
11
+ 2026-03-22 23:43:10,252 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
12
+ 2026-03-22 23:43:19,148 [INFO] ==> starting training
13
+ 2026-03-22 23:43:19,150 [INFO] ==> loading model: /app
14
+ 2026-03-22 23:43:19,151 [INFO] ==> logical base model id: Qwen/Qwen2.5-7B-Instruct
15
+ 2026-03-22 23:43:19,151 [INFO] ==> method: qlora
16
+ 2026-03-22 23:43:19,152 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
17
+ 2026-03-22 23:43:19,459 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
18
+ 2026-03-22 23:43:19,942 [INFO] HTTP Request: GET https://huggingface.co/api/whoami-v2 "HTTP/1.1 200 OK"
19
+ 2026-03-22 23:43:28,508 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
20
+ 2026-03-22 23:43:29,000 [INFO] HTTP Request: HEAD https://s3.amazonaws.com/datasets.huggingface.co/datasets/datasets/json/json.py "HTTP/1.1 200 OK"
21
+ 2026-03-22 23:43:32,143 [INFO] Unsloth: Padding-free batching auto-enabled for SFTTrainer instance.
22
+ 2026-03-22 23:43:32,334 [WARNING] num_proc must be <= 2. Reducing num_proc to 2 for dataset of size 2.
23
+ 2026-03-22 23:43:37,654 [WARNING] num_proc must be <= 1. Reducing num_proc to 1 for dataset of size 1.
24
+ 2026-03-22 23:43:41,594 [INFO] ==> training started
25
+ 2026-03-22 23:43:42,839 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
26
+ 2026-03-22 23:43:50,075 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
27
+ 2026-03-22 23:43:50,729 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
28
+ 2026-03-22 23:43:50,963 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
29
+ 2026-03-22 23:43:51,219 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
30
+ 2026-03-22 23:43:51,348 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
31
+ 2026-03-22 23:43:51,568 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
32
+ 2026-03-22 23:43:51,580 [INFO] ==> training finished
33
+ 2026-03-22 23:43:51,580 [INFO] ==> saving lora adapters to /output/job_e2e_1774222987861/lora/trainer-e2e-job_e2e_1774222987861
34
+ 2026-03-22 23:43:51,595 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
35
+ 2026-03-22 23:43:51,612 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
36
+ 2026-03-22 23:43:51,630 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
37
+ 2026-03-22 23:43:51,766 [INFO] ==> saving merged model to /output/job_e2e_1774222987861/merged/trainer-e2e-job_e2e_1774222987861
38
+ 2026-03-22 23:43:51,766 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
39
+ 2026-03-22 23:43:51,766 [INFO] ==> merged 16-bit save may take several minutes for 7B model
40
+ 2026-03-22 23:44:16,075 [INFO] ==> merged model saved
41
+ 2026-03-22 23:44:16,079 [INFO] ==> training artifacts saved
42
+ 2026-03-22 23:44:16,079 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
43
+ 2026-03-22 23:44:16,079 [INFO] ==> training finished
44
+ 2026-03-22 23:44:16,088 [INFO] ==> starting evaluation
45
+ 2026-03-22 23:44:16,132 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
46
+ 2026-03-22 23:44:22,174 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
47
+ 2026-03-22 23:44:22,232 [INFO] ==> reporting progress to http://host.docker.internal:18787/api/jobs/progress
48
+ 2026-03-22 23:44:22,356 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
49
+ 2026-03-22 23:44:22,357 [INFO] ==> evaluation finished
50
+ 2026-03-22 23:44:22,403 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
51
+ 2026-03-22 23:44:22,469 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar from /output/job_e2e_1774222987861/lora/trainer-e2e-job_e2e_1774222987861 (mode=w, files=6, size_bytes=21538633)
52
+ 2026-03-22 23:44:22,500 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar
53
+ 2026-03-22 23:44:22,500 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/lora: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar.gz
54
+ 2026-03-22 23:44:22,502 [ERROR] upload step failed: lora_archive
55
  Traceback (most recent call last):
56
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
57
  result = operation()
 
64
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
65
  with open(archive_path, "rb") as f:
66
  ^^^^^^^^^^^^^^^^^^^^^^^^
67
+ FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.lora.tar.gz'
68
+ 2026-03-22 23:44:22,504 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar from /output/job_e2e_1774222987861/merged/trainer-e2e-job_e2e_1774222987861 (mode=w, files=9, size_bytes=15242728923)
69
+ 2026-03-22 23:44:33,444 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar
70
+ 2026-03-22 23:44:33,444 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/merged: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar.gz
71
+ 2026-03-22 23:44:33,445 [ERROR] upload step failed: merged_archive
72
  Traceback (most recent call last):
73
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
74
  result = operation()
 
81
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
82
  with open(archive_path, "rb") as f:
83
  ^^^^^^^^^^^^^^^^^^^^^^^^
84
+ FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.merged.tar.gz'
85
+ 2026-03-22 23:44:33,449 [INFO] ==> creating archive /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar from /output/job_e2e_1774222987861 (mode=w, files=36, size_bytes=15291110697)
86
+ 2026-03-22 23:44:44,091 [INFO] ==> archive ready: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar
87
+ 2026-03-22 23:44:44,091 [INFO] ==> uploading archive to http://host.docker.internal:18787/api/jobs/upload/full-archive: /output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar.gz
88
+ 2026-03-22 23:44:44,091 [ERROR] upload step failed: full_archive
89
  Traceback (most recent call last):
90
  File "/trainer/app/pipeline/upload_runner.py", line 85, in _safe_upload
91
  result = operation()
 
98
  File "/trainer/app/pipeline/archiver.py", line 190, in upload_archive
99
  with open(archive_path, "rb") as f:
100
  ^^^^^^^^^^^^^^^^^^^^^^^^
101
+ FileNotFoundError: [Errno 2] No such file or directory: '/output/job_e2e_1774222987861/trainer-e2e-job_e2e_1774222987861.full.tar.gz'
102
+ 2026-03-22 23:44:44,136 [INFO] ==> reporting status to http://host.docker.internal:18787/api/jobs/status
103
+ 2026-03-22 23:44:44,391 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
104
+ 2026-03-22 23:45:21,012 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"
105
+ 2026-03-22 23:45:21,184 [INFO] HTTP Request: POST https://huggingface.co/XProger/test.git/info/lfs/objects/batch "HTTP/1.1 200 OK"
106
+ 2026-03-22 23:45:21,438 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/xet-write-token/main "HTTP/1.1 200 OK"
107
+ 2026-03-22 23:45:48,946 [WARNING] No files have been modified since last commit. Skipping to prevent empty commit.
108
+ 2026-03-22 23:45:49,203 [INFO] HTTP Request: GET https://huggingface.co/api/models/XProger/test/revision/main "HTTP/1.1 200 OK"
109
+ 2026-03-22 23:45:49,384 [INFO] HTTP Request: POST https://huggingface.co/api/repos/create "HTTP/1.1 409 Conflict"
110
+ 2026-03-22 23:45:49,567 [INFO] HTTP Request: POST https://huggingface.co/api/models/XProger/test/preupload/main "HTTP/1.1 200 OK"