stefan-it's picture
Upload folder using huggingface_hub
9343f28
2023-10-17 18:57:52,091 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,093 Model: "SequenceTagger(
(embeddings): TransformerWordEmbeddings(
(model): ElectraModel(
(embeddings): ElectraEmbeddings(
(word_embeddings): Embedding(32001, 768)
(position_embeddings): Embedding(512, 768)
(token_type_embeddings): Embedding(2, 768)
(LayerNorm): LayerNorm((768,), eps=1e-12, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
(encoder): ElectraEncoder(
(layer): ModuleList(
(0-11): 12 x ElectraLayer(
(attention): ElectraAttention(
(self): ElectraSelfAttention(
(query): Linear(in_features=768, out_features=768, bias=True)
(key): Linear(in_features=768, out_features=768, bias=True)
(value): Linear(in_features=768, out_features=768, bias=True)
(dropout): Dropout(p=0.1, inplace=False)
)
(output): ElectraSelfOutput(
(dense): Linear(in_features=768, out_features=768, bias=True)
(LayerNorm): LayerNorm((768,), eps=1e-12, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
)
(intermediate): ElectraIntermediate(
(dense): Linear(in_features=768, out_features=3072, bias=True)
(intermediate_act_fn): GELUActivation()
)
(output): ElectraOutput(
(dense): Linear(in_features=3072, out_features=768, bias=True)
(LayerNorm): LayerNorm((768,), eps=1e-12, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
)
)
)
)
)
(locked_dropout): LockedDropout(p=0.5)
(linear): Linear(in_features=768, out_features=21, bias=True)
(loss_function): CrossEntropyLoss()
)"
2023-10-17 18:57:52,093 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,093 MultiCorpus: 3575 train + 1235 dev + 1266 test sentences
- NER_HIPE_2022 Corpus: 3575 train + 1235 dev + 1266 test sentences - /root/.flair/datasets/ner_hipe_2022/v2.1/hipe2020/de/with_doc_seperator
2023-10-17 18:57:52,093 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,093 Train: 3575 sentences
2023-10-17 18:57:52,094 (train_with_dev=False, train_with_test=False)
2023-10-17 18:57:52,094 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,094 Training Params:
2023-10-17 18:57:52,094 - learning_rate: "5e-05"
2023-10-17 18:57:52,094 - mini_batch_size: "4"
2023-10-17 18:57:52,094 - max_epochs: "10"
2023-10-17 18:57:52,094 - shuffle: "True"
2023-10-17 18:57:52,094 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,094 Plugins:
2023-10-17 18:57:52,094 - TensorboardLogger
2023-10-17 18:57:52,094 - LinearScheduler | warmup_fraction: '0.1'
2023-10-17 18:57:52,094 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,094 Final evaluation on model from best epoch (best-model.pt)
2023-10-17 18:57:52,094 - metric: "('micro avg', 'f1-score')"
2023-10-17 18:57:52,094 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,095 Computation:
2023-10-17 18:57:52,095 - compute on device: cuda:0
2023-10-17 18:57:52,095 - embedding storage: none
2023-10-17 18:57:52,095 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,095 Model training base path: "hmbench-hipe2020/de-hmteams/teams-base-historic-multilingual-discriminator-bs4-wsFalse-e10-lr5e-05-poolingfirst-layers-1-crfFalse-5"
2023-10-17 18:57:52,095 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,095 ----------------------------------------------------------------------------------------------------
2023-10-17 18:57:52,095 Logging anything other than scalars to TensorBoard is currently not supported.
2023-10-17 18:57:58,931 epoch 1 - iter 89/894 - loss 3.00225299 - time (sec): 6.83 - samples/sec: 1202.01 - lr: 0.000005 - momentum: 0.000000
2023-10-17 18:58:05,870 epoch 1 - iter 178/894 - loss 1.79069375 - time (sec): 13.77 - samples/sec: 1206.03 - lr: 0.000010 - momentum: 0.000000
2023-10-17 18:58:13,106 epoch 1 - iter 267/894 - loss 1.31705555 - time (sec): 21.01 - samples/sec: 1210.50 - lr: 0.000015 - momentum: 0.000000
2023-10-17 18:58:19,938 epoch 1 - iter 356/894 - loss 1.07438081 - time (sec): 27.84 - samples/sec: 1201.72 - lr: 0.000020 - momentum: 0.000000
2023-10-17 18:58:26,897 epoch 1 - iter 445/894 - loss 0.91287612 - time (sec): 34.80 - samples/sec: 1213.00 - lr: 0.000025 - momentum: 0.000000
2023-10-17 18:58:33,882 epoch 1 - iter 534/894 - loss 0.82028036 - time (sec): 41.79 - samples/sec: 1201.28 - lr: 0.000030 - momentum: 0.000000
2023-10-17 18:58:41,226 epoch 1 - iter 623/894 - loss 0.73685559 - time (sec): 49.13 - samples/sec: 1193.86 - lr: 0.000035 - momentum: 0.000000
2023-10-17 18:58:48,391 epoch 1 - iter 712/894 - loss 0.66554276 - time (sec): 56.29 - samples/sec: 1203.17 - lr: 0.000040 - momentum: 0.000000
2023-10-17 18:58:55,466 epoch 1 - iter 801/894 - loss 0.60697439 - time (sec): 63.37 - samples/sec: 1212.86 - lr: 0.000045 - momentum: 0.000000
2023-10-17 18:59:03,146 epoch 1 - iter 890/894 - loss 0.56625986 - time (sec): 71.05 - samples/sec: 1213.13 - lr: 0.000050 - momentum: 0.000000
2023-10-17 18:59:03,472 ----------------------------------------------------------------------------------------------------
2023-10-17 18:59:03,472 EPOCH 1 done: loss 0.5641 - lr: 0.000050
2023-10-17 18:59:10,298 DEV : loss 0.1641440987586975 - f1-score (micro avg) 0.6191
2023-10-17 18:59:10,363 saving best model
2023-10-17 18:59:10,899 ----------------------------------------------------------------------------------------------------
2023-10-17 18:59:18,143 epoch 2 - iter 89/894 - loss 0.15794860 - time (sec): 7.24 - samples/sec: 1382.36 - lr: 0.000049 - momentum: 0.000000
2023-10-17 18:59:25,658 epoch 2 - iter 178/894 - loss 0.17315256 - time (sec): 14.76 - samples/sec: 1259.52 - lr: 0.000049 - momentum: 0.000000
2023-10-17 18:59:32,636 epoch 2 - iter 267/894 - loss 0.16955826 - time (sec): 21.73 - samples/sec: 1218.46 - lr: 0.000048 - momentum: 0.000000
2023-10-17 18:59:39,880 epoch 2 - iter 356/894 - loss 0.16310974 - time (sec): 28.98 - samples/sec: 1213.18 - lr: 0.000048 - momentum: 0.000000
2023-10-17 18:59:47,181 epoch 2 - iter 445/894 - loss 0.15719978 - time (sec): 36.28 - samples/sec: 1193.70 - lr: 0.000047 - momentum: 0.000000
2023-10-17 18:59:54,203 epoch 2 - iter 534/894 - loss 0.15641959 - time (sec): 43.30 - samples/sec: 1216.41 - lr: 0.000047 - momentum: 0.000000
2023-10-17 19:00:01,701 epoch 2 - iter 623/894 - loss 0.15556446 - time (sec): 50.80 - samples/sec: 1197.75 - lr: 0.000046 - momentum: 0.000000
2023-10-17 19:00:08,829 epoch 2 - iter 712/894 - loss 0.15363448 - time (sec): 57.93 - samples/sec: 1199.92 - lr: 0.000046 - momentum: 0.000000
2023-10-17 19:00:16,271 epoch 2 - iter 801/894 - loss 0.15304683 - time (sec): 65.37 - samples/sec: 1187.87 - lr: 0.000045 - momentum: 0.000000
2023-10-17 19:00:23,793 epoch 2 - iter 890/894 - loss 0.15066106 - time (sec): 72.89 - samples/sec: 1180.58 - lr: 0.000044 - momentum: 0.000000
2023-10-17 19:00:24,111 ----------------------------------------------------------------------------------------------------
2023-10-17 19:00:24,111 EPOCH 2 done: loss 0.1504 - lr: 0.000044
2023-10-17 19:00:35,210 DEV : loss 0.18480995297431946 - f1-score (micro avg) 0.6824
2023-10-17 19:00:35,280 saving best model
2023-10-17 19:00:36,773 ----------------------------------------------------------------------------------------------------
2023-10-17 19:00:44,207 epoch 3 - iter 89/894 - loss 0.10262602 - time (sec): 7.43 - samples/sec: 1156.20 - lr: 0.000044 - momentum: 0.000000
2023-10-17 19:00:51,543 epoch 3 - iter 178/894 - loss 0.10632283 - time (sec): 14.77 - samples/sec: 1129.54 - lr: 0.000043 - momentum: 0.000000
2023-10-17 19:00:59,146 epoch 3 - iter 267/894 - loss 0.10131406 - time (sec): 22.37 - samples/sec: 1159.43 - lr: 0.000043 - momentum: 0.000000
2023-10-17 19:01:06,517 epoch 3 - iter 356/894 - loss 0.10708151 - time (sec): 29.74 - samples/sec: 1154.81 - lr: 0.000042 - momentum: 0.000000
2023-10-17 19:01:13,599 epoch 3 - iter 445/894 - loss 0.10588775 - time (sec): 36.82 - samples/sec: 1169.24 - lr: 0.000042 - momentum: 0.000000
2023-10-17 19:01:20,881 epoch 3 - iter 534/894 - loss 0.10796682 - time (sec): 44.10 - samples/sec: 1193.24 - lr: 0.000041 - momentum: 0.000000
2023-10-17 19:01:27,787 epoch 3 - iter 623/894 - loss 0.10768458 - time (sec): 51.01 - samples/sec: 1189.83 - lr: 0.000041 - momentum: 0.000000
2023-10-17 19:01:34,808 epoch 3 - iter 712/894 - loss 0.10710551 - time (sec): 58.03 - samples/sec: 1190.19 - lr: 0.000040 - momentum: 0.000000
2023-10-17 19:01:41,848 epoch 3 - iter 801/894 - loss 0.10295354 - time (sec): 65.07 - samples/sec: 1188.45 - lr: 0.000039 - momentum: 0.000000
2023-10-17 19:01:49,373 epoch 3 - iter 890/894 - loss 0.10328161 - time (sec): 72.60 - samples/sec: 1186.01 - lr: 0.000039 - momentum: 0.000000
2023-10-17 19:01:49,715 ----------------------------------------------------------------------------------------------------
2023-10-17 19:01:49,715 EPOCH 3 done: loss 0.1031 - lr: 0.000039
2023-10-17 19:02:00,859 DEV : loss 0.1962013840675354 - f1-score (micro avg) 0.7387
2023-10-17 19:02:00,916 saving best model
2023-10-17 19:02:02,326 ----------------------------------------------------------------------------------------------------
2023-10-17 19:02:09,542 epoch 4 - iter 89/894 - loss 0.04452472 - time (sec): 7.21 - samples/sec: 1048.57 - lr: 0.000038 - momentum: 0.000000
2023-10-17 19:02:16,679 epoch 4 - iter 178/894 - loss 0.05905823 - time (sec): 14.35 - samples/sec: 1153.82 - lr: 0.000038 - momentum: 0.000000
2023-10-17 19:02:23,887 epoch 4 - iter 267/894 - loss 0.06476324 - time (sec): 21.56 - samples/sec: 1202.03 - lr: 0.000037 - momentum: 0.000000
2023-10-17 19:02:30,893 epoch 4 - iter 356/894 - loss 0.07137867 - time (sec): 28.56 - samples/sec: 1221.77 - lr: 0.000037 - momentum: 0.000000
2023-10-17 19:02:38,140 epoch 4 - iter 445/894 - loss 0.07551925 - time (sec): 35.81 - samples/sec: 1207.72 - lr: 0.000036 - momentum: 0.000000
2023-10-17 19:02:45,733 epoch 4 - iter 534/894 - loss 0.07291703 - time (sec): 43.40 - samples/sec: 1197.24 - lr: 0.000036 - momentum: 0.000000
2023-10-17 19:02:53,181 epoch 4 - iter 623/894 - loss 0.07411859 - time (sec): 50.85 - samples/sec: 1195.67 - lr: 0.000035 - momentum: 0.000000
2023-10-17 19:03:00,307 epoch 4 - iter 712/894 - loss 0.07308605 - time (sec): 57.98 - samples/sec: 1194.32 - lr: 0.000034 - momentum: 0.000000
2023-10-17 19:03:07,370 epoch 4 - iter 801/894 - loss 0.07206437 - time (sec): 65.04 - samples/sec: 1201.71 - lr: 0.000034 - momentum: 0.000000
2023-10-17 19:03:14,207 epoch 4 - iter 890/894 - loss 0.07099225 - time (sec): 71.88 - samples/sec: 1199.37 - lr: 0.000033 - momentum: 0.000000
2023-10-17 19:03:14,508 ----------------------------------------------------------------------------------------------------
2023-10-17 19:03:14,508 EPOCH 4 done: loss 0.0707 - lr: 0.000033
2023-10-17 19:03:25,597 DEV : loss 0.20233656466007233 - f1-score (micro avg) 0.7493
2023-10-17 19:03:25,659 saving best model
2023-10-17 19:03:27,156 ----------------------------------------------------------------------------------------------------
2023-10-17 19:03:34,492 epoch 5 - iter 89/894 - loss 0.03267879 - time (sec): 7.33 - samples/sec: 1167.64 - lr: 0.000033 - momentum: 0.000000
2023-10-17 19:03:41,819 epoch 5 - iter 178/894 - loss 0.04292784 - time (sec): 14.66 - samples/sec: 1177.46 - lr: 0.000032 - momentum: 0.000000
2023-10-17 19:03:49,250 epoch 5 - iter 267/894 - loss 0.03836055 - time (sec): 22.09 - samples/sec: 1210.20 - lr: 0.000032 - momentum: 0.000000
2023-10-17 19:03:56,478 epoch 5 - iter 356/894 - loss 0.03997784 - time (sec): 29.32 - samples/sec: 1197.60 - lr: 0.000031 - momentum: 0.000000
2023-10-17 19:04:03,597 epoch 5 - iter 445/894 - loss 0.04386553 - time (sec): 36.44 - samples/sec: 1178.16 - lr: 0.000031 - momentum: 0.000000
2023-10-17 19:04:10,959 epoch 5 - iter 534/894 - loss 0.04250446 - time (sec): 43.80 - samples/sec: 1178.17 - lr: 0.000030 - momentum: 0.000000
2023-10-17 19:04:18,388 epoch 5 - iter 623/894 - loss 0.04348624 - time (sec): 51.23 - samples/sec: 1177.13 - lr: 0.000029 - momentum: 0.000000
2023-10-17 19:04:25,406 epoch 5 - iter 712/894 - loss 0.04515069 - time (sec): 58.25 - samples/sec: 1172.83 - lr: 0.000029 - momentum: 0.000000
2023-10-17 19:04:33,093 epoch 5 - iter 801/894 - loss 0.04710607 - time (sec): 65.93 - samples/sec: 1159.01 - lr: 0.000028 - momentum: 0.000000
2023-10-17 19:04:40,554 epoch 5 - iter 890/894 - loss 0.04592844 - time (sec): 73.40 - samples/sec: 1174.34 - lr: 0.000028 - momentum: 0.000000
2023-10-17 19:04:40,881 ----------------------------------------------------------------------------------------------------
2023-10-17 19:04:40,881 EPOCH 5 done: loss 0.0457 - lr: 0.000028
2023-10-17 19:04:52,692 DEV : loss 0.2360163778066635 - f1-score (micro avg) 0.7779
2023-10-17 19:04:52,750 saving best model
2023-10-17 19:04:54,180 ----------------------------------------------------------------------------------------------------
2023-10-17 19:05:01,845 epoch 6 - iter 89/894 - loss 0.01149976 - time (sec): 7.66 - samples/sec: 1263.07 - lr: 0.000027 - momentum: 0.000000
2023-10-17 19:05:09,303 epoch 6 - iter 178/894 - loss 0.02451550 - time (sec): 15.12 - samples/sec: 1179.73 - lr: 0.000027 - momentum: 0.000000
2023-10-17 19:05:16,170 epoch 6 - iter 267/894 - loss 0.02732072 - time (sec): 21.99 - samples/sec: 1179.96 - lr: 0.000026 - momentum: 0.000000
2023-10-17 19:05:23,068 epoch 6 - iter 356/894 - loss 0.02567363 - time (sec): 28.88 - samples/sec: 1186.14 - lr: 0.000026 - momentum: 0.000000
2023-10-17 19:05:30,144 epoch 6 - iter 445/894 - loss 0.02565507 - time (sec): 35.96 - samples/sec: 1181.82 - lr: 0.000025 - momentum: 0.000000
2023-10-17 19:05:37,584 epoch 6 - iter 534/894 - loss 0.02724422 - time (sec): 43.40 - samples/sec: 1183.74 - lr: 0.000024 - momentum: 0.000000
2023-10-17 19:05:44,541 epoch 6 - iter 623/894 - loss 0.02748912 - time (sec): 50.36 - samples/sec: 1189.75 - lr: 0.000024 - momentum: 0.000000
2023-10-17 19:05:51,801 epoch 6 - iter 712/894 - loss 0.02961830 - time (sec): 57.62 - samples/sec: 1205.02 - lr: 0.000023 - momentum: 0.000000
2023-10-17 19:05:59,381 epoch 6 - iter 801/894 - loss 0.03061449 - time (sec): 65.20 - samples/sec: 1200.70 - lr: 0.000023 - momentum: 0.000000
2023-10-17 19:06:06,705 epoch 6 - iter 890/894 - loss 0.03059783 - time (sec): 72.52 - samples/sec: 1187.95 - lr: 0.000022 - momentum: 0.000000
2023-10-17 19:06:07,039 ----------------------------------------------------------------------------------------------------
2023-10-17 19:06:07,041 EPOCH 6 done: loss 0.0306 - lr: 0.000022
2023-10-17 19:06:18,622 DEV : loss 0.26993367075920105 - f1-score (micro avg) 0.7554
2023-10-17 19:06:18,677 ----------------------------------------------------------------------------------------------------
2023-10-17 19:06:26,150 epoch 7 - iter 89/894 - loss 0.01061319 - time (sec): 7.47 - samples/sec: 1180.43 - lr: 0.000022 - momentum: 0.000000
2023-10-17 19:06:33,151 epoch 7 - iter 178/894 - loss 0.01720578 - time (sec): 14.47 - samples/sec: 1185.25 - lr: 0.000021 - momentum: 0.000000
2023-10-17 19:06:40,480 epoch 7 - iter 267/894 - loss 0.01736450 - time (sec): 21.80 - samples/sec: 1167.73 - lr: 0.000021 - momentum: 0.000000
2023-10-17 19:06:47,520 epoch 7 - iter 356/894 - loss 0.01641373 - time (sec): 28.84 - samples/sec: 1182.45 - lr: 0.000020 - momentum: 0.000000
2023-10-17 19:06:54,381 epoch 7 - iter 445/894 - loss 0.01756040 - time (sec): 35.70 - samples/sec: 1190.16 - lr: 0.000019 - momentum: 0.000000
2023-10-17 19:07:01,235 epoch 7 - iter 534/894 - loss 0.01848064 - time (sec): 42.56 - samples/sec: 1192.17 - lr: 0.000019 - momentum: 0.000000
2023-10-17 19:07:08,148 epoch 7 - iter 623/894 - loss 0.01888162 - time (sec): 49.47 - samples/sec: 1201.31 - lr: 0.000018 - momentum: 0.000000
2023-10-17 19:07:15,100 epoch 7 - iter 712/894 - loss 0.01793327 - time (sec): 56.42 - samples/sec: 1200.36 - lr: 0.000018 - momentum: 0.000000
2023-10-17 19:07:22,436 epoch 7 - iter 801/894 - loss 0.01765536 - time (sec): 63.76 - samples/sec: 1220.72 - lr: 0.000017 - momentum: 0.000000
2023-10-17 19:07:29,336 epoch 7 - iter 890/894 - loss 0.01794690 - time (sec): 70.66 - samples/sec: 1218.98 - lr: 0.000017 - momentum: 0.000000
2023-10-17 19:07:29,666 ----------------------------------------------------------------------------------------------------
2023-10-17 19:07:29,667 EPOCH 7 done: loss 0.0179 - lr: 0.000017
2023-10-17 19:07:41,350 DEV : loss 0.2559952437877655 - f1-score (micro avg) 0.7682
2023-10-17 19:07:41,415 ----------------------------------------------------------------------------------------------------
2023-10-17 19:07:48,666 epoch 8 - iter 89/894 - loss 0.00418217 - time (sec): 7.25 - samples/sec: 1098.61 - lr: 0.000016 - momentum: 0.000000
2023-10-17 19:07:55,712 epoch 8 - iter 178/894 - loss 0.00926357 - time (sec): 14.29 - samples/sec: 1168.98 - lr: 0.000016 - momentum: 0.000000
2023-10-17 19:08:02,788 epoch 8 - iter 267/894 - loss 0.01003390 - time (sec): 21.37 - samples/sec: 1267.76 - lr: 0.000015 - momentum: 0.000000
2023-10-17 19:08:09,602 epoch 8 - iter 356/894 - loss 0.01001150 - time (sec): 28.18 - samples/sec: 1243.38 - lr: 0.000014 - momentum: 0.000000
2023-10-17 19:08:16,385 epoch 8 - iter 445/894 - loss 0.01084819 - time (sec): 34.97 - samples/sec: 1239.62 - lr: 0.000014 - momentum: 0.000000
2023-10-17 19:08:23,255 epoch 8 - iter 534/894 - loss 0.01123504 - time (sec): 41.84 - samples/sec: 1242.38 - lr: 0.000013 - momentum: 0.000000
2023-10-17 19:08:31,006 epoch 8 - iter 623/894 - loss 0.01130336 - time (sec): 49.59 - samples/sec: 1226.23 - lr: 0.000013 - momentum: 0.000000
2023-10-17 19:08:38,228 epoch 8 - iter 712/894 - loss 0.01157568 - time (sec): 56.81 - samples/sec: 1219.93 - lr: 0.000012 - momentum: 0.000000
2023-10-17 19:08:45,047 epoch 8 - iter 801/894 - loss 0.01127330 - time (sec): 63.63 - samples/sec: 1210.31 - lr: 0.000012 - momentum: 0.000000
2023-10-17 19:08:52,150 epoch 8 - iter 890/894 - loss 0.01057093 - time (sec): 70.73 - samples/sec: 1219.24 - lr: 0.000011 - momentum: 0.000000
2023-10-17 19:08:52,453 ----------------------------------------------------------------------------------------------------
2023-10-17 19:08:52,453 EPOCH 8 done: loss 0.0105 - lr: 0.000011
2023-10-17 19:09:03,092 DEV : loss 0.27024897933006287 - f1-score (micro avg) 0.7726
2023-10-17 19:09:03,152 ----------------------------------------------------------------------------------------------------
2023-10-17 19:09:10,821 epoch 9 - iter 89/894 - loss 0.00360434 - time (sec): 7.67 - samples/sec: 1040.93 - lr: 0.000011 - momentum: 0.000000
2023-10-17 19:09:18,205 epoch 9 - iter 178/894 - loss 0.00565799 - time (sec): 15.05 - samples/sec: 1090.50 - lr: 0.000010 - momentum: 0.000000
2023-10-17 19:09:25,863 epoch 9 - iter 267/894 - loss 0.00507366 - time (sec): 22.71 - samples/sec: 1097.93 - lr: 0.000009 - momentum: 0.000000
2023-10-17 19:09:33,952 epoch 9 - iter 356/894 - loss 0.00531567 - time (sec): 30.80 - samples/sec: 1121.02 - lr: 0.000009 - momentum: 0.000000
2023-10-17 19:09:41,522 epoch 9 - iter 445/894 - loss 0.00513708 - time (sec): 38.37 - samples/sec: 1101.26 - lr: 0.000008 - momentum: 0.000000
2023-10-17 19:09:48,777 epoch 9 - iter 534/894 - loss 0.00571170 - time (sec): 45.62 - samples/sec: 1104.59 - lr: 0.000008 - momentum: 0.000000
2023-10-17 19:09:56,160 epoch 9 - iter 623/894 - loss 0.00624185 - time (sec): 53.01 - samples/sec: 1111.21 - lr: 0.000007 - momentum: 0.000000
2023-10-17 19:10:03,766 epoch 9 - iter 712/894 - loss 0.00585911 - time (sec): 60.61 - samples/sec: 1125.45 - lr: 0.000007 - momentum: 0.000000
2023-10-17 19:10:11,350 epoch 9 - iter 801/894 - loss 0.00569284 - time (sec): 68.20 - samples/sec: 1138.08 - lr: 0.000006 - momentum: 0.000000
2023-10-17 19:10:18,597 epoch 9 - iter 890/894 - loss 0.00557183 - time (sec): 75.44 - samples/sec: 1142.99 - lr: 0.000006 - momentum: 0.000000
2023-10-17 19:10:18,924 ----------------------------------------------------------------------------------------------------
2023-10-17 19:10:18,924 EPOCH 9 done: loss 0.0056 - lr: 0.000006
2023-10-17 19:10:30,263 DEV : loss 0.2831907570362091 - f1-score (micro avg) 0.7883
2023-10-17 19:10:30,335 saving best model
2023-10-17 19:10:31,798 ----------------------------------------------------------------------------------------------------
2023-10-17 19:10:39,599 epoch 10 - iter 89/894 - loss 0.00821683 - time (sec): 7.80 - samples/sec: 1176.09 - lr: 0.000005 - momentum: 0.000000
2023-10-17 19:10:46,872 epoch 10 - iter 178/894 - loss 0.00610798 - time (sec): 15.07 - samples/sec: 1267.45 - lr: 0.000004 - momentum: 0.000000
2023-10-17 19:10:54,031 epoch 10 - iter 267/894 - loss 0.00562470 - time (sec): 22.23 - samples/sec: 1240.04 - lr: 0.000004 - momentum: 0.000000
2023-10-17 19:11:01,068 epoch 10 - iter 356/894 - loss 0.00514968 - time (sec): 29.27 - samples/sec: 1240.12 - lr: 0.000003 - momentum: 0.000000
2023-10-17 19:11:07,994 epoch 10 - iter 445/894 - loss 0.00468347 - time (sec): 36.19 - samples/sec: 1229.33 - lr: 0.000003 - momentum: 0.000000
2023-10-17 19:11:15,362 epoch 10 - iter 534/894 - loss 0.00476647 - time (sec): 43.56 - samples/sec: 1215.10 - lr: 0.000002 - momentum: 0.000000
2023-10-17 19:11:22,828 epoch 10 - iter 623/894 - loss 0.00468070 - time (sec): 51.03 - samples/sec: 1196.79 - lr: 0.000002 - momentum: 0.000000
2023-10-17 19:11:30,064 epoch 10 - iter 712/894 - loss 0.00443002 - time (sec): 58.26 - samples/sec: 1186.03 - lr: 0.000001 - momentum: 0.000000
2023-10-17 19:11:37,851 epoch 10 - iter 801/894 - loss 0.00419002 - time (sec): 66.05 - samples/sec: 1172.63 - lr: 0.000001 - momentum: 0.000000
2023-10-17 19:11:44,539 epoch 10 - iter 890/894 - loss 0.00386993 - time (sec): 72.74 - samples/sec: 1182.57 - lr: 0.000000 - momentum: 0.000000
2023-10-17 19:11:44,847 ----------------------------------------------------------------------------------------------------
2023-10-17 19:11:44,847 EPOCH 10 done: loss 0.0038 - lr: 0.000000
2023-10-17 19:11:56,138 DEV : loss 0.2942058742046356 - f1-score (micro avg) 0.7877
2023-10-17 19:11:56,763 ----------------------------------------------------------------------------------------------------
2023-10-17 19:11:56,765 Loading model from best epoch ...
2023-10-17 19:11:59,022 SequenceTagger predicts: Dictionary with 21 tags: O, S-loc, B-loc, E-loc, I-loc, S-pers, B-pers, E-pers, I-pers, S-org, B-org, E-org, I-org, S-prod, B-prod, E-prod, I-prod, S-time, B-time, E-time, I-time
2023-10-17 19:12:04,771
Results:
- F-score (micro) 0.76
- F-score (macro) 0.6777
- Accuracy 0.6269
By class:
precision recall f1-score support
loc 0.8608 0.8406 0.8506 596
pers 0.7180 0.7417 0.7297 333
org 0.5577 0.4394 0.4915 132
prod 0.6875 0.5000 0.5789 66
time 0.7037 0.7755 0.7379 49
micro avg 0.7747 0.7457 0.7600 1176
macro avg 0.7055 0.6594 0.6777 1176
weighted avg 0.7701 0.7457 0.7561 1176
2023-10-17 19:12:04,771 ----------------------------------------------------------------------------------------------------