Couverture de la campagne
60/89 exécutées
17 FAIL, 4 neutralisées et 29 non exécutées, présentées séparément.
Dernière campagne : 39 tâches réussies sur les 89 officielles avec DeepSeek V4 Flash, soit 43,8 %. Retrouvez les 60 tâches exécutées, les 29 non exécutées, les cas neutralisés et la consommation mesurée, sans confondre des modèles ou protocoles différents.
DeepSeek V4 Flash via Souver Desktop : 43,8 % sur les 89 tâches officielles, avec 39 PASS, 17 FAIL et 4 mesures neutralisées pour un incident d’infrastructure du vérificateur. 29 tâches n’ont pas été exécutées. Les neutralisations et les non-exécutions restent distinctes des échecs du modèle ; le dénominateur du score reste 89.
Couverture de la campagne
60/89 exécutées
17 FAIL, 4 neutralisées et 29 non exécutées, présentées séparément.
Septembre · DeepSeek V4 Flash
39/89 · 43,8 %
39 PASS. Aucun résultat n’est imputé aux tâches non exécutées.
Coût opérationnel conservateur
45,91 €
45,908 € incluant API et infrastructure. Estimation, pas facture débitée.
Desktop 9922ff0a8c54b76fa53489ff68adcf5b3a25491d · App 167a0c21c92ba830db5e84d51400708b001cc53a · Harness f81872b413b83f65f4a0b7025ff2baa5554d3945
Source tb60-terminal-results.json · SHA256 61f08d56b84c4e97e76530aca6d81074ebc5594b2c9bbc4da8f2c19f17d2b3f4
Télécharger les 60 résultats et leurs mesures (JSON)| Tâche | Verdict | Durée cumulée | Requêtes | Crédits | Entrée / cache / sortie |
|---|---|---|---|---|---|
| fix-git | PASS | 119 s | 19 | 18 547 | 174 677 / 116 224 / 5 379 |
| prove-plus-comm | PASS | 165 s | 38 | 39 513 | 423 890 / 307 712 / 9 626 |
| cobol-modernization | PASS | 193 s | 73 | 135 047 | 1 504 475 / 1 094 144 / 22 103 |
| overfull-hbox | FAIL | 320 s | 31 | 77 101 | 611 054 / 334 592 / 20 541 |
| crack-7z-hash | PASS | 313 s | 31 | 35 688 | 325 385 / 194 816 / 4 195 |
| mteb-leaderboard | FAIL | 625 s | 9 | 6 914 | 60 344 / 41 216 / 3 543 |
| raman-fitting | FAIL | 295 s | 57 | 131 994 | 1 304 460 / 889 856 / 32 799 |
| constraints-scheduling | PASS | 122 s | 7 | 14 123 | 86 258 / 42 496 / 9 082 |
| kv-store-grpc | PASS | 87 s | 27 | 28 412 | 242 250 / 139 264 / 5 391 |
| mteb-retrieve | FAIL | 695 s | 12 | 10 390 | 88 305 / 52 992 / 2 936 |
| pytorch-model-recovery | PASS | 819 s | 17 | 36 233 | 211 780 / 60 160 / 8 508 |
| break-filter-js-from-html | PASS | 544 s | 27 | 76 184 | 580 058 / 313 856 / 25 462 |
| hf-model-inference | PASS | 406 s | 22 | 20 594 | 202 260 / 136 960 / 4 983 |
| merge-diff-arc-agi-task | PASS | 191 s | 40 | 151 224 | 1 256 099 / 661 504 / 13 614 |
| nginx-request-logging | PASS | 134 s | 29 | 39 868 | 387 613 / 250 624 / 5 829 |
| openssl-selfsigned-cert | PASS | 97 s | 22 | 19 502 | 161 956 / 91 136 / 4 089 |
| polyglot-c-py | FAIL | 181 s | 14 | 51 971 | 317 359 / 129 536 / 22 719 |
| vulnerable-secret | PASS | 96 s | 18 | 26 793 | 211 582 / 111 360 / 5 548 |
| code-from-image | FAIL | 951 s | 110 | 1 165 920 | 9 947 232 / 5 362 688 / 78 821 |
| count-dataset-tokens | FAIL | 158 s | 33 | 43 637 | 446 264 / 315 392 / 11 806 |
| custom-memory-heap-crash | PASS | 282 s | 43 | 102 921 | 964 398 / 619 008 / 22 005 |
| dna-insert | FAIL | 433 s | 59 | 269 815 | 2 625 425 / 1 758 208 / 63 347 |
| extract-elf | FAIL | 265 s | 26 | 247 170 | 1 794 679 / 765 696 / 25 284 |
| financial-document-processor | PASS | 323 s | 38 | 95 707 | 861 449 / 509 184 / 11 574 |
| git-leak-recovery | PASS | 77 s | 16 | 10 827 | 113 761 / 82 944 / 3 272 |
| multi-source-data-merger | PASS | 129 s | 15 | 19 166 | 151 946 / 87 040 / 6 621 |
| pytorch-model-cli | PASS | 361 s | 30 | 42 385 | 370 187 / 222 464 / 9 548 |
| qemu-alpine-ssh | Neutralisée infra-verificateur-sentinelle | 918 s | 58 | 158 107 | 1 248 663 / 649 728 / 29 762 |
| qemu-startup | Neutralisée infra-verificateur-sentinelle | 1 001 s | 79 | 519 725 | 4 699 745 / 2 701 568 / 26 723 |
| reshard-c4-data | PASS | 403 s | 23 | 298 299 | 2 121 755 / 844 032 / 20 588 |
| sanitize-git-repo | FAIL | 621 s | 114 | 793 435 | 10 467 611 / 8 307 200 / 67 564 |
| sqlite-with-gcov | PASS | 364 s | 58 | 286 330 | 2 629 712 / 1 519 616 / 6 958 |
| tune-mjcf | PASS | 942 s | 51 | 167 574 | 1 186 562 / 500 480 / 24 731 |
| large-scale-text-editing | PASS | 238 s | 17 | 29 831 | 234 236 / 129 792 / 9 170 |
| chess-best-move | FAIL | 398 s | 38 | 233 590 | 1 893 021 / 1 070 080 / 63 985 |
| db-wal-recovery | PASS | 128 s | 28 | 87 256 | 669 700 / 314 368 / 8 455 |
| filter-js-from-html | Neutralisée infra-verificateur-sentinelle | 379 s | 11 | 11 040 | 89 932 / 57 344 / 5 485 |
| regex-log | PASS | 126 s | 13 | 27 765 | 187 895 / 93 440 / 12 656 |
| build-cython-ext | PASS | 409 s | 102 | 300 765 | 3 492 710 / 2 539 776 / 19 455 |
| build-pov-ray | FAIL | 669 s | 94 | 954 374 | 9 346 297 / 5 817 344 / 33 639 |
| compile-compcert | PASS | 1 606 s | 64 | 468 022 | 3 643 643 / 1 672 704 / 14 306 |
| gcode-to-text | FAIL | 919 s | 95 | 932 191 | 9 592 106 / 6 411 520 / 93 082 |
| largest-eigenval | PASS | 829 s | 75 | 283 315 | 2 784 657 / 1 826 304 / 44 607 |
| mailman | PASS | 428 s | 62 | 166 811 | 1 765 822 / 1 226 240 / 23 488 |
| pypi-server | PASS | 246 s | 35 | 24 898 | 296 101 / 228 352 / 5 332 |
| query-optimize | Neutralisée infra-verificateur-sentinelle | 1 205 s | 18 | 33 850 | 272 749 / 140 032 / 4 030 |
| sqlite-db-truncate | PASS | 281 s | 14 | 50 070 | 259 363 / 119 808 / 43 088 |
| winning-avg-corewars | PASS | 754 s | 62 | 343 810 | 2 827 266 / 1 554 432 / 65 439 |
| log-summary-date-ranges | FAIL | 199 s | 20 | 36 288 | 280 565 / 139 008 / 5 792 |
| build-pmars | PASS | 478 s | 44 | 156 922 | 1 464 468 / 873 472 / 8 422 |
| distribution-search | PASS | 136 s | 25 | 57 802 | 446 296 / 238 592 / 16 396 |
| headless-terminal | PASS | 232 s | 26 | 37 419 | 303 388 / 174 080 / 11 219 |
| modernize-scientific-stack | PASS | 135 s | 11 | 8 285 | 82 624 / 58 368 / 2 681 |
| portfolio-optimization | PASS | 501 s | 34 | 84 291 | 664 269 / 357 120 / 20 867 |
| adaptive-rejection-sampler | FAIL | 727 s | 22 | 144 939 | 957 505 / 382 720 / 35 745 |
| git-multibranch | PASS | 362 s | 46 | 67 372 | 722 059 / 515 584 / 13 156 |
| rstan-to-pystan | FAIL | 1 828 s | 63 | 279 557 | 2 284 226 / 1 171 968 / 23 734 |
| schemelike-metacircular-eval | PASS | 1 133 s | 97 | 548 078 | 6 910 941 / 5 444 352 / 88 911 |
| caffe-cifar-10 | FAIL | 1 305 s | 28 | 145 177 | 1 232 326 / 652 544 / 6 853 |
| configure-git-webserver | PASS | 353 s | 77 | 280 428 | 3 540 033 / 2 725 888 / 19 556 |
Aucun verdict mesuré : elles ne sont ni des FAIL ni des neutralisations.
Chaque carte pointe vers une cohorte homogène. Aucun shard plus récent ne remplace silencieusement un full 89 tâches.
GPT-5.5
72.7 %
Score · 64 PASS sur 88 résultats comptables
Run full-tb2-k1-runner-1-runner-1-s0-20260711T222729Z-d0a2efc3 · Desktop 0.7.151
GLM 5.2
43.2 %
Score · 38 PASS sur 88 résultats comptables
Run full-tb2-k1-20260703T131756Z · Desktop 0.7.130
DeepSeek V4 Pro
34.1 %
Score · 30 PASS sur 88 résultats comptables
Run full-tb2-k1-20260706T083738Z · Desktop 0.7.145
Chaque cellule conserve son état propre. Une panne provider ou un essai non exécuté n'est jamais transformé en échec du modèle. Cliquez sur une cellule d'évaluation pour ouvrir son détail audité.
89 tâches affichées
| Tâche TB2 | GPT-5.5 | DeepSeek V4 Flash39/89 · 43,8 % | GLM 5.2 | DeepSeek V4 Pro |
|---|---|---|---|---|
| adaptive-rejection-sampler | ||||
| bn-fit-modify | ||||
| break-filter-js-from-html | ||||
| build-cython-ext | ||||
| build-pmars | ||||
| build-pov-ray | ||||
| caffe-cifar-10 | ||||
| cancel-async-tasks | ||||
| chess-best-move | ||||
| circuit-fibsqrt | ||||
| cobol-modernization | ||||
| code-from-image | ||||
| compile-compcert | ||||
| configure-git-webserver | ||||
| constraints-scheduling | ||||
| count-dataset-tokens | ||||
| crack-7z-hash | ||||
| custom-memory-heap-crash | ||||
| db-wal-recovery | ||||
| distribution-search | ||||
| dna-assembly | ||||
| dna-insert | ||||
| extract-elf | ||||
| extract-moves-from-video | ||||
| feal-differential-cryptanalysis | ||||
| feal-linear-cryptanalysis | ||||
| filter-js-from-html | ||||
| financial-document-processor | ||||
| fix-code-vulnerability | ||||
| fix-git | ||||
| fix-ocaml-gc | ||||
| gcode-to-text | ||||
| git-leak-recovery | ||||
| git-multibranch | ||||
| gpt2-codegolf | ||||
| headless-terminal | ||||
| hf-model-inference | ||||
| install-windows-3.11 | ||||
| kv-store-grpc | ||||
| large-scale-text-editing | ||||
| largest-eigenval | ||||
| llm-inference-batching-scheduler | ||||
| log-summary-date-ranges | ||||
| mailman | ||||
| make-doom-for-mips | ||||
| make-mips-interpreter | ||||
| mcmc-sampling-stan | ||||
| merge-diff-arc-agi-task | ||||
| model-extraction-relu-logits | ||||
| modernize-scientific-stack | ||||
| mteb-leaderboard | ||||
| mteb-retrieve | ||||
| multi-source-data-merger | ||||
| nginx-request-logging | ||||
| openssl-selfsigned-cert | ||||
| overfull-hbox | ||||
| password-recovery | ||||
| path-tracing | ||||
| path-tracing-reverse | ||||
| polyglot-c-py | ||||
| polyglot-rust-c | ||||
| portfolio-optimization | ||||
| protein-assembly | ||||
| prove-plus-comm | ||||
| pypi-server | ||||
| pytorch-model-cli | ||||
| pytorch-model-recovery | ||||
| qemu-alpine-ssh | ||||
| qemu-startup | ||||
| query-optimize | ||||
| raman-fitting | ||||
| regex-chess | ||||
| regex-log | ||||
| reshard-c4-data | ||||
| rstan-to-pystan | ||||
| sam-cell-seg | ||||
| sanitize-git-repo | ||||
| schemelike-metacircular-eval | ||||
| sparql-university | ||||
| sqlite-db-truncate | ||||
| sqlite-with-gcov | ||||
| torch-pipeline-parallelism | ||||
| torch-tensor-parallelism | ||||
| train-fasttext | ||||
| tune-mjcf | ||||
| video-processing | ||||
| vulnerable-secret | ||||
| winning-avg-corewars | ||||
| write-compressor |
Échangeons sur vos cas d'usage, vos modèles et vos contraintes : nous vous aidons à construire un harness plus fiable, plus efficient et adapté à votre environnement.
Planifier un échange de 30 min →