| 1 | fd9bd632f346085920d523dd307a3e4c11cc0a05 | e08e473cf292ba73789769b993bb35bcdfc31689 | wangbomeng | wangbomeng@calculet.tech | 2026-06-01T14:15:35+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-06-01T14:15:35+08:00 | HEAD -> dev, origin/dev | llama-bench: fix device-info |
|---|
| 2 | e08e473cf292ba73789769b993bb35bcdfc31689 | f86112d6758298241ab2fa84ad4f0f7b1ace35c6 | wangbomeng | wangbomeng@calculet.tech | 2026-06-01T11:47:01+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-06-01T11:47:01+08:00 | | llama-bench:add device-info |
| 3 | f86112d6758298241ab2fa84ad4f0f7b1ace35c6 | d1bc743c10080fed1253686ce01749b42611ded4 | wangbomeng | wangbomeng@calculet.tech | 2026-05-29T16:40:38+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-29T16:40:38+08:00 | | add dump_pos,fix pos_tensor of pre_by_embeds |
| 4 | d1bc743c10080fed1253686ce01749b42611ded4 | 2f201160b1960fd7be18b70d5770599d7082dd2f | wangbomeng | wangbomeng@calculet.tech | 2026-05-28T14:23:06+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-28T14:23:06+08:00 | | set u_batch = max_seq_len |
| 5 | 2f201160b1960fd7be18b70d5770599d7082dd2f | f131bf1908b8c5dd1f49d7ae7f616506fc471666 | wangbomeng | wangbomeng@calculet.tech | 2026-05-27T09:58:31+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-27T09:58:31+08:00 | | delete useless code |
| 6 | f131bf1908b8c5dd1f49d7ae7f616506fc471666 | 33d96dae28e17208ef61a71dba40cda3c04f135a | wangbomeng | wangbomeng@calculet.tech | 2026-05-27T09:38:51+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-27T09:38:51+08:00 | | fix patch of qwen2.5vl,n_patch of minicpm = 64,fix n_ctx_slot,comment out M-ROPE |
| 7 | ebce9e1a548c2329aa97f53b8b2ad50d968088dc | e850ff4de25ad1221abc9bf3bb30df252ccd9c36 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-27T09:38:00+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-27T09:38:00+08:00 | origin/runtime_replace | Refactor cal-llm integration by removing runtime capacity loading and adjusting request handling |
| 8 | e850ff4de25ad1221abc9bf3bb30df252ccd9c36 | 9985688f7fa96f755882f96a2cd8a18114fdc42d | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-26T10:59:45+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-26T10:59:45+08:00 | | Enhance cal-llm integration by adding runtime checks for CALRT_LIBRARY and updating link settings for Linux and Windows |
| 9 | 9985688f7fa96f755882f96a2cd8a18114fdc42d | 38ca6895815671a87ceec86b4aa4eb976580ce76 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-25T09:05:49+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-25T09:05:49+08:00 | | Refactor CalrtEngineConfig initialization by removing redundant parameters |
| 10 | 33d96dae28e17208ef61a71dba40cda3c04f135a | 37f068d87cc2d0a0b40ce978031fc9932fc89077 | wangbomeng | wangbomeng@calculet.tech | 2026-05-22T14:33:45+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-22T14:33:45+08:00 | | add llama-build.sh |
| 11 | 1921024604db640f09ade83c8a437ebf15bde360 | d028722a12c5a463cc97207787cfd03d7a2590b0 5af229fe518ee2d761ab8ae2a69f7a9aaa9057f4 | Bomeng Wang | wangbomeng@calculet.tech | 2026-05-21T14:08:14+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-05-21T14:08:14+08:00 | refs/stash | WIP on dev: d028722a replace cparam.n_ubatch with n_ubatch() |
| 12 | 5af229fe518ee2d761ab8ae2a69f7a9aaa9057f4 | d028722a12c5a463cc97207787cfd03d7a2590b0 | Bomeng Wang | wangbomeng@calculet.tech | 2026-05-21T14:08:14+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-05-21T14:08:14+08:00 | | index on dev: d028722a replace cparam.n_ubatch with n_ubatch() |
| 13 | 37f068d87cc2d0a0b40ce978031fc9932fc89077 | 13508e5e18bbb1202ae60c2ee5ecbdda1cca817b | wangbomeng | wangbomeng@calculet.tech | 2026-05-21T14:07:00+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-21T14:07:00+08:00 | | release |
| 14 | 13508e5e18bbb1202ae60c2ee5ecbdda1cca817b | b6b7919468126d5ba30a10941825154789e9082f | wangbomeng | wangbomeng@calculet.tech | 2026-05-21T11:43:59+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-21T11:43:59+08:00 | | reduce cmake log |
| 15 | b6b7919468126d5ba30a10941825154789e9082f | 46847c24fc5c52dd57bb4fb17fc4d684149bb4fa | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T16:40:33+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T16:40:33+08:00 | | reduce cmake log |
| 16 | 38ca6895815671a87ceec86b4aa4eb976580ce76 | ddd60c928224b52eb4e1e60e01e8a1ecbb80ee45 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-20T16:37:13+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-20T16:37:13+08:00 | | Add support for cal-llm profile dumping to JSONL file |
| 17 | 46847c24fc5c52dd57bb4fb17fc4d684149bb4fa | d028722a12c5a463cc97207787cfd03d7a2590b0 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T16:02:12+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T16:02:12+08:00 | | Simplify CMakeLists.txt |
| 18 | d028722a12c5a463cc97207787cfd03d7a2590b0 | 63442fde8dc285817f725bd6c5fae80f478a3813 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T14:41:05+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T14:41:05+08:00 | | replace cparam.n_ubatch with n_ubatch() |
| 19 | 63442fde8dc285817f725bd6c5fae80f478a3813 | fc3f4a9fab88e98eff8eeca8cbc6617333edba2c | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T11:47:42+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-20T11:47:42+08:00 | | fix ubatch and cmakelist |
| 20 | ddd60c928224b52eb4e1e60e01e8a1ecbb80ee45 | 2716130010c58f1cdd9d94f0e2f8e047bc92dc66 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-20T09:22:40+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-20T09:27:06+08:00 | | Add cal-llm backend profiling support with command-line option |
| 21 | fc3f4a9fab88e98eff8eeca8cbc6617333edba2c | c931ef3163940227c9b6291e04216245cce21429 | wangbomeng | wangbomeng@calculet.tech | 2026-05-19T16:54:16+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-19T16:54:16+08:00 | | add dump and rt case generation |
| 22 | c931ef3163940227c9b6291e04216245cce21429 | a43741244a646d3f8fa3ff01bb9b7d059223d0a5 | wangbomeng | wangbomeng@calculet.tech | 2026-05-18T17:34:37+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-18T17:34:37+08:00 | | calrt: fix prefill_by_ids.fix create_buf |
| 23 | 2716130010c58f1cdd9d94f0e2f8e047bc92dc66 | b35bda124ccd4ed581dd6088d3ffa4931c2809b6 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-18T08:40:54+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-18T08:40:54+08:00 | | fix n_ctx_slot error by adding runtime capacity loading for cal-llm |
| 24 | a43741244a646d3f8fa3ff01bb9b7d059223d0a5 | b2003d6c90ef0fa26ab9e930ba3e22b355eb3d39 | wangbomeng | wangbomeng@calculet.tech | 2026-05-15T14:39:53+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-15T14:39:53+08:00 | | run llama-passkey successfully, fix llama-bench, copy 3 times input to infer qwen2.5vl |
| 25 | b35bda124ccd4ed581dd6088d3ffa4931c2809b6 | 81a2cbd4cef5d28334ff513d6b69e593d9d1cb49 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-15T09:19:31+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-15T09:19:31+08:00 | | Implement vocab-only model initialization and enhance cal-llm backend error handling |
| 26 | 81a2cbd4cef5d28334ff513d6b69e593d9d1cb49 | 3b14ed556c15529fe1f94b649f49db7156dfaf67 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-14T16:06:37+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-14T16:06:37+08:00 | | Enhance CMake configuration for cal-llm integration by updating paths and adding library properties |
| 27 | 3b14ed556c15529fe1f94b649f49db7156dfaf67 | dde8b8228081769c8d8f1e0751d47fa7875f8423 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-14T14:34:25+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-05-14T14:34:25+08:00 | | - Removed references to CALRT from clip.h, mtmd-helper.cpp, mtmd-helper.h, mtmd.cpp, and mtmd.h. - Updated server.cpp to conditionally include CAL-LLM support. - Introduced option LLAMA_SERVER_USE_CAL_LLM in CMakeLists.txt to enable CAL-LLM integration. - Refactored server context to handle CAL-LLM requests and responses. - Updated image processing logic to remove CALRT-specific implementations. - Ensured compatibility with existing llama functionality while integrating new CAL-LLM features. |
| 28 | b2003d6c90ef0fa26ab9e930ba3e22b355eb3d39 | 7800c772d44a7415640ec1669a691232939927cf | wangbomeng | wangbomeng@calculet.tech | 2026-05-13T15:24:24+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-13T15:24:24+08:00 | | try passkey and embedding |
| 29 | 7800c772d44a7415640ec1669a691232939927cf | dde8b8228081769c8d8f1e0751d47fa7875f8423 | wangbomeng | wangbomeng@calculet.tech | 2026-05-13T15:23:54+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-13T15:23:54+08:00 | | try passkey and embedding |
| 30 | dde8b8228081769c8d8f1e0751d47fa7875f8423 | fabcc7dac4b34d83fe430980bb364b81087c7ee9 | wangbomeng | wangbomeng@calculet.tech | 2026-05-12T10:46:08+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-05-12T10:46:08+08:00 | | clip:add minicpm patch; |
| 31 | fabcc7dac4b34d83fe430980bb364b81087c7ee9 | 1e9787962f20a7e82c71589ac1dd0585454e4bfe | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:17:17+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:17:17+08:00 | | calrt: fix encode dump |
| 32 | 1e9787962f20a7e82c71589ac1dd0585454e4bfe | f584311c716dd078c4b737eebbfda97fe12b7274 1bec0db5a675ddc60fb793be2a746e8f3c8855dc | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:06:23+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:06:23+08:00 | | Merge remote-tracking branch 'origin/dev-ben' into dev |
| 33 | f584311c716dd078c4b737eebbfda97fe12b7274 | a339954cc2a145a80c5ea349e7896211916cbed9 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:05:15+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:05:15+08:00 | | calrt: fix encode dump |
| 34 | 1bec0db5a675ddc60fb793be2a746e8f3c8855dc | 465df20c3e1f3f58d0f5531c796e0941a49c32b0 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:02:59+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T16:02:59+08:00 | origin/dev-ben | bench: support 2 calbins |
| 35 | 465df20c3e1f3f58d0f5531c796e0941a49c32b0 | a339954cc2a145a80c5ea349e7896211916cbed9 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T09:08:12+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-30T09:08:12+08:00 | | support one calbin |
| 36 | a339954cc2a145a80c5ea349e7896211916cbed9 | 8e76fc330ad624d1768b412d7894e272dbe696f4 eeeef3f6d6a5aad9296323ca15d9d929745b8932 | wangbomeng | wangbomeng@calculet.tech | 2026-04-29T14:57:55+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-29T14:57:55+08:00 | tag: v0.3.0 | Merge branch 'dev' into dev-cli |
| 37 | 8e76fc330ad624d1768b412d7894e272dbe696f4 | 431ad6f1c787f1b47e9e7be14b2e1a7d8565485c | wangbomeng | wangbomeng@calculet.tech | 2026-04-29T14:56:59+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-29T14:56:59+08:00 | | finish dev of llama-cli |
| 38 | eeeef3f6d6a5aad9296323ca15d9d929745b8932 | 431ad6f1c787f1b47e9e7be14b2e1a7d8565485c | wangbomeng | wangbomeng@calculet.tech | 2026-04-28T16:16:19+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-28T16:16:19+08:00 | | add unknown time info: prefill and decode |
| 39 | 431ad6f1c787f1b47e9e7be14b2e1a7d8565485c | efa1876bb8b2a449a55054a8c3745892f2c0f262 | wangbomeng | wangbomeng@calculet.tech | 2026-04-27T16:52:15+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-27T16:52:15+08:00 | | update calrt_infer |
| 40 | efa1876bb8b2a449a55054a8c3745892f2c0f262 | 4a2a3181f1cea170105c771e6ba69d851b82b937 | wangbomeng | wangbomeng@calculet.tech | 2026-04-27T11:35:58+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-27T11:35:58+08:00 | tag: v0.2.2 | update version print |
| 41 | 4a2a3181f1cea170105c771e6ba69d851b82b937 | e3ea5541259f35871c3e995331770aaca297e9ec | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T10:56:02+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T10:56:02+08:00 | | add calrt-version |
| 42 | e3ea5541259f35871c3e995331770aaca297e9ec | 0b8e43943846a1cd05c26c9186b9ef778ebae6e2 | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T10:55:48+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T10:55:48+08:00 | | add calrt-version |
| 43 | 0b8e43943846a1cd05c26c9186b9ef778ebae6e2 | 5eb46fffbd2e260c540fb112314ad67bbb8f56b6 | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T08:43:43+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-26T08:43:43+08:00 | | to debug |
| 44 | 5eb46fffbd2e260c540fb112314ad67bbb8f56b6 | 005112c51e00d998b457337ca3dad53fc46e2d0b | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T16:24:51+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T16:24:51+08:00 | | clip: add image_to_patches; calrt: vision copy data from bf16 to fp32 |
| 45 | 005112c51e00d998b457337ca3dad53fc46e2d0b | 509670642e8090a1713f5d73590b549be827aaef | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T16:24:31+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T16:24:31+08:00 | | clip: add image_to_patches; calrt: vision copy data from bf16 to fp32 |
| 46 | 509670642e8090a1713f5d73590b549be827aaef | 5d219bfcef239c1e3ddd1baecc0514a9462ff6a8 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:40:34+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:40:34+08:00 | | rm debug code |
| 47 | 5d219bfcef239c1e3ddd1baecc0514a9462ff6a8 | 6787a53ab9b4897c5a7b51a4062d7b90e6697d35 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:31:32+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:31:32+08:00 | | calrt: recover base = batch_index * dict_len |
| 48 | 6787a53ab9b4897c5a7b51a4062d7b90e6697d35 | e1a76abf66178ea85656479a5e19aea7785c09f4 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:22:02+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-24T09:22:02+08:00 | | calrt:input of multimodel from fp32 to bf16 |
| 49 | e1a76abf66178ea85656479a5e19aea7785c09f4 | 99dd308adb8488fa869bb669e3335e646a938eff | wangbomeng | wangbomeng@calculet.tech | 2026-04-22T16:29:22+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-22T16:29:22+08:00 | | calrt:add CalcoreRT version print |
| 50 | 99dd308adb8488fa869bb669e3335e646a938eff | 06c7852e99e34c71e24e3ef6001cb54c72970e4d | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:50:01+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:50:01+08:00 | tag: v0.2.1 | server: fix max_seq_len |
| 51 | 06c7852e99e34c71e24e3ef6001cb54c72970e4d | 96b4e267e562f44e76ec6c965ee0fa361f0439bf ed62eb33447f993d578e156a3c7e38c91b420ece | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:06:01+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:06:01+08:00 | | Merge branch 'dev-ot' into dev |
| 52 | ed62eb33447f993d578e156a3c7e38c91b420ece | 015e988a271573a7c11e2a4b45084e1219d69d50 | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:05:22+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-21T15:05:22+08:00 | | try server-test:not success |
| 53 | 015e988a271573a7c11e2a4b45084e1219d69d50 | 99962021f2dd1994dbbf90d5712f0ca2633f20f8 | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T16:56:08+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T16:56:08+08:00 | | little change |
| 54 | 96b4e267e562f44e76ec6c965ee0fa361f0439bf | 99962021f2dd1994dbbf90d5712f0ca2633f20f8 | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T15:33:01+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T15:33:01+08:00 | | merge dev-ot |
| 55 | 99962021f2dd1994dbbf90d5712f0ca2633f20f8 | 9af6682ef101c6d006c743c703cf150ac25f466f | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T14:48:33+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-17T14:48:33+08:00 | origin/dev-ot | calrt: add onetoken slice and untile_prefill |
| 56 | 9af6682ef101c6d006c743c703cf150ac25f466f | e24f775ee26ea5249d5ff6cd23b8a0a17fe1438d | wangbomeng | wangbomeng@calculet.tech | 2026-04-16T10:09:44+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-16T10:09:44+08:00 | | args: add -ca |
| 57 | e24f775ee26ea5249d5ff6cd23b8a0a17fe1438d | a206460e610aef995e36e50b33e8bd41fe477541 | wangbomeng | wangbomeng@calculet.tech | 2026-04-15T10:01:09+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-15T10:01:09+08:00 | | calrt: add copy_data_to_ibuf for vision model |
| 58 | 795179beb5bba728e0ccf0a7e80a5ec5ef36e92a | 6d55eae7e76b768acfcec045b8e84cded8ae3b55 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-14T16:03:22+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-14T16:03:22+08:00 | origin/test-cicd | add ci yaml |
| 59 | a206460e610aef995e36e50b33e8bd41fe477541 | d8b2aaa00049978cebac15041f960d75b552f97b | wangbomeng | wangbomeng@calculet.tech | 2026-04-14T10:23:43+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-14T10:23:43+08:00 | | calrt: support mixture of prefill_by_ids and prefill_by_embeds |
| 60 | d8b2aaa00049978cebac15041f960d75b552f97b | b6f29ec8140028619efaa50aa6e73146cecf631d | wangbomeng | wangbomeng@calculet.tech | 2026-04-13T11:23:44+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-13T11:23:44+08:00 | | server: fix context-shift |
| 61 | b6f29ec8140028619efaa50aa6e73146cecf631d | 6d55eae7e76b768acfcec045b8e84cded8ae3b55 | wangbomeng | wangbomeng@calculet.tech | 2026-04-13T11:23:31+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-13T11:23:31+08:00 | | server: fix context-shift |
| 62 | 6d55eae7e76b768acfcec045b8e84cded8ae3b55 | af8055f5f033a730e7e1fe80185b18b7e9473966 | wangbomeng | wangbomeng@calculet.tech | 2026-04-10T14:58:56+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-10T14:58:56+08:00 | tag: v0.2.0 | add annotations |
| 63 | af8055f5f033a730e7e1fe80185b18b7e9473966 | 6f54682df62b8ec6bb4154b99c47b5becda3291f | wangbomeng | wangbomeng@calculet.tech | 2026-04-10T14:58:19+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-10T14:58:19+08:00 | | add annotations |
| 64 | 6f54682df62b8ec6bb4154b99c47b5becda3291f | 42a181674565411efa5c8b318bc77d6fd084df8a a2d5bf362d6c9a546f9deadf4f8975605a915d33 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-10T09:07:20+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-10T09:07:20+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 65 | 42a181674565411efa5c8b318bc77d6fd084df8a | bd3a57a11aa008e0b5dc2a30f82fcd3a02587208 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-10T09:06:55+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-10T09:06:55+08:00 | | fix prompt_cache issue |
| 66 | a2d5bf362d6c9a546f9deadf4f8975605a915d33 | bd3a57a11aa008e0b5dc2a30f82fcd3a02587208 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-09T21:01:46+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-09T21:01:46+08:00 | | server: comment fixed time |
| 67 | bd3a57a11aa008e0b5dc2a30f82fcd3a02587208 | fcfcdf80b01664a08fb12df879a05df3cc0eba70 74d437ab1456831573def7ca385a74d0ca460c37 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T19:42:57+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T19:42:57+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 68 | fcfcdf80b01664a08fb12df879a05df3cc0eba70 | 0cd44929b442860dc0e5f88976108be82350f9dc | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T19:41:52+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T19:41:52+08:00 | | tmp fix one run >4K problem |
| 69 | 74d437ab1456831573def7ca385a74d0ca460c37 | 5813c943005d1ede551e67ee98e35aefd210ff22 0cd44929b442860dc0e5f88976108be82350f9dc | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:44:26+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:44:26+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 70 | 5813c943005d1ede551e67ee98e35aefd210ff22 | 1f0d5bb0453f93ffa51faf68237c89251c2e0c40 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:38:57+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:38:57+08:00 | | add --device-info |
| 71 | 1f0d5bb0453f93ffa51faf68237c89251c2e0c40 | b8153b6b8f244b223c9f4c2940ae1f212634519e | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:38:37+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T17:38:37+08:00 | | add --device-info |
| 72 | 0cd44929b442860dc0e5f88976108be82350f9dc | b8153b6b8f244b223c9f4c2940ae1f212634519e | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-09T15:08:50+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-09T15:08:50+08:00 | | rm calculet0.txt |
| 73 | b8153b6b8f244b223c9f4c2940ae1f212634519e | 82842cfc3fa3d5c004c4540aeff8a81f38509bee | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T15:07:02+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T15:07:02+08:00 | | server: fix limit tokens to max_seq_len |
| 74 | 82842cfc3fa3d5c004c4540aeff8a81f38509bee | 622a22991089c6debd5f4b57738f3df3a0bda8b5 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T15:06:26+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T15:06:26+08:00 | | server: fix limit tokens to max_seq_len |
| 75 | 622a22991089c6debd5f4b57738f3df3a0bda8b5 | 5f47a738ba56a37d0ff281a16212f41979568e35 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T10:36:31+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-09T10:36:31+08:00 | | server: comment circle print |
| 76 | 5f47a738ba56a37d0ff281a16212f41979568e35 | ecf01a17651cd7b1e8b85fd6075e83831463ee23 207ca38c751f7e6c251833eb11e89a2a0bb0fd07 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T09:48:20+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T09:48:20+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 77 | ecf01a17651cd7b1e8b85fd6075e83831463ee23 | 99d2078e89305035d87d966cee7987c7d50b3925 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T09:47:55+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-09T09:47:55+08:00 | | fix n_parallel problem |
| 78 | 207ca38c751f7e6c251833eb11e89a2a0bb0fd07 | 99d2078e89305035d87d966cee7987c7d50b3925 | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T19:14:37+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T19:14:37+08:00 | | delet extra printf |
| 79 | 99d2078e89305035d87d966cee7987c7d50b3925 | 9626239f5a9e568c9975d67dd6745fef90279a87 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-08T19:01:16+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-04-08T19:01:16+08:00 | | fix buffer resize problem |
| 80 | 9626239f5a9e568c9975d67dd6745fef90279a87 | da724f3e2a8af2c2537e3edcff41e9415ac66fef | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T17:47:04+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T17:47:04+08:00 | | calrt: fix MapBuf |
| 81 | da724f3e2a8af2c2537e3edcff41e9415ac66fef | 18814f9bf1a423ebe8f32d6043c9cbc5b21b0218 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-08T16:58:32+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-08T16:58:32+08:00 | | calrt: add MapBuf |
| 82 | 18814f9bf1a423ebe8f32d6043c9cbc5b21b0218 | bbee06d6d851e3f84f75b578047967d7e0e4a8dc | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T13:55:52+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T13:55:52+08:00 | | server : reset inference time |
| 83 | bbee06d6d851e3f84f75b578047967d7e0e4a8dc | e126c21ed7b793b648d63533ea7ee30885f5a82a | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T11:29:36+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-08T11:29:36+08:00 | | calrt: add test time print |
| 84 | 39edc6ec2e9560ccb20a2b8e547f58102bb268c1 | 2006831fd6ea8a8e16518b81bcee29fb674676ea | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-01T14:06:50+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-04-01T14:06:50+08:00 | tag: v0.1.1 | add t_incopy t_outcopy |
| 85 | 2006831fd6ea8a8e16518b81bcee29fb674676ea | 41e6636ea1fb3884e41fb2023f92b0e7262ef09b | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:42:27+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:42:27+08:00 | tag: v0.1.0 | v0.1.0 |
| 86 | e126c21ed7b793b648d63533ea7ee30885f5a82a | 5603e050af07173589f3b8850121263d86da01fd | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:07:12+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:07:12+08:00 | | sever: fix commit id |
| 87 | 5603e050af07173589f3b8850121263d86da01fd | 1e8a16cd0e3ad3c9a0f8746852a0a38c6e52322e | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T05:57:55+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T05:57:55+08:00 | | calrt: add t_incopy and t_outcopy |
| 88 | 1e8a16cd0e3ad3c9a0f8746852a0a38c6e52322e | 893e52bda0fee99bbda1521ea733a760bc4201f1 | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:14:22+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-04-01T10:14:22+08:00 | | calrt: fix slice and add infer/copy time |
| 89 | 893e52bda0fee99bbda1521ea733a760bc4201f1 | 398ecf71f5f574037ce33f84dd1576fa164909ad | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-30T11:35:40+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-30T11:35:40+08:00 | | delete unnecessary commit id |
| 90 | 398ecf71f5f574037ce33f84dd1576fa164909ad | 6d5e83902bdd3142a8ff0d3db19f02a8619e95a0 | wangbomeng | wangbomeng@calculet.tech | 2026-03-27T02:01:23+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-27T02:01:23+08:00 | | rm llama-server_old.cpp |
| 91 | 6d5e83902bdd3142a8ff0d3db19f02a8619e95a0 | e4eaab700333490e80ebefc8d023f6c1adfcc312 | wangbomeng | wangbomeng@calculet.tech | 2026-03-27T00:57:11+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-27T00:57:11+08:00 | | server: printf calrt commit ID |
| 92 | e4eaab700333490e80ebefc8d023f6c1adfcc312 | aeacc23a883400db0ebffb50e23355ada054da66 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T09:03:55+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T09:03:55+08:00 | | calrt: prefil to prefill |
| 93 | aeacc23a883400db0ebffb50e23355ada054da66 | cbaa7492663630132adeddccd9fec5e44ffa3336 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T08:49:22+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T08:49:22+08:00 | | calrt: add slice only on decode |
| 94 | cbaa7492663630132adeddccd9fec5e44ffa3336 | 3c08ebf0e442cd730ddffd959ec676f9caf31c82 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T03:21:02+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-25T03:21:02+08:00 | | reduce warning |
| 95 | 3c08ebf0e442cd730ddffd959ec676f9caf31c82 | 613bc2b9a9d7af92efbe42d876ffcdb244f7a00e | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T02:43:03+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T02:43:03+08:00 | | calrt: comment LOG_ERROR in untile_decode_cpy |
| 96 | 613bc2b9a9d7af92efbe42d876ffcdb244f7a00e | 4e194ecc6c4b8ebe4c72d93f6055e3605dbfcefa | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T02:30:12+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T02:30:12+08:00 | | calrt: fix max_seq_len judgment |
| 97 | 4e194ecc6c4b8ebe4c72d93f6055e3605dbfcefa | 16cbf81a731f5839a2f6b0eb9d8539826353748b | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T01:59:03+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-24T01:59:03+08:00 | | calrt: fix ubatch.embd cpy |
| 98 | 16cbf81a731f5839a2f6b0eb9d8539826353748b | 8b4e17d990c23025148c4abadefe1010a0dd0734 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T09:38:21+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T09:38:21+08:00 | | calrt: fix model_name of untile_cpy |
| 99 | 8b4e17d990c23025148c4abadefe1010a0dd0734 | 6dacdd2cf9574214ca272a86674e7bd7ba21c4f0 41e6636ea1fb3884e41fb2023f92b0e7262ef09b | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:50:27+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:50:27+08:00 | | merge origin dev |
| 100 | 6dacdd2cf9574214ca272a86674e7bd7ba21c4f0 | 3e6ad64467771224f072e5bad90e270951ba4cdf | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:42:28+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:42:28+08:00 | | calrt: add multimodal |
| 101 | 3e6ad64467771224f072e5bad90e270951ba4cdf | b37b7d9756feb3dac2db526eeab38b572a095d47 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:36:41+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-23T14:36:41+08:00 | | calrt: add multimodal |
| 102 | 41e6636ea1fb3884e41fb2023f92b0e7262ef09b | d0a1b2c758440720d42fa108b3444f54593daa73 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-12T16:36:07+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-12T16:36:07+08:00 | tag: v0.0.1 | server: correct the max_seq_len judgment |
| 103 | d0a1b2c758440720d42fa108b3444f54593daa73 | b37b7d9756feb3dac2db526eeab38b572a095d47 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-12T16:10:28+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-12T16:10:28+08:00 | | calrt: comment out dump |
| 104 | b37b7d9756feb3dac2db526eeab38b572a095d47 | 07817cefc14fa359dcecad0630defd3610f8f5c8 7ddafbdcb9426ae3af016925b8b596f11e8742d0 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-10T16:32:31+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-10T16:32:31+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 105 | 07817cefc14fa359dcecad0630defd3610f8f5c8 | b2c45d8ddefb7ec09435231c8ec5b736fabf9c91 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-10T16:32:07+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-10T16:32:07+08:00 | | fix kv issue |
| 106 | 7ddafbdcb9426ae3af016925b8b596f11e8742d0 | 059009eeaf20a569abfbac5eb37dad3bb9376428 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T07:13:04+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T07:13:04+08:00 | | calrt: commented out cal_ctx->encode(batch) |
| 107 | 059009eeaf20a569abfbac5eb37dad3bb9376428 | 1b073661d1311b5300173993b9f31ee7241d8606 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T07:04:22+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T07:04:22+08:00 | | calrt: add multimodel |
| 108 | 1b073661d1311b5300173993b9f31ee7241d8606 | 6295273c5866b9249bacdedbfc3a31877e6e6a85 b2c45d8ddefb7ec09435231c8ec5b736fabf9c91 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T06:58:35+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T06:58:35+08:00 | | Merge branch 'dev' of http://192.168.20.229:1000/yunzhe/llama.cpp into dev |
| 109 | 6295273c5866b9249bacdedbfc3a31877e6e6a85 | c6ea3f0d2468e3b8eab5de80bacee7fe4c9dd3d7 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T06:54:41+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-10T06:54:41+08:00 | | arg: update hf_repo url to https://download.calculet.tech:9443 |
| 110 | b2c45d8ddefb7ec09435231c8ec5b736fabf9c91 | c6ea3f0d2468e3b8eab5de80bacee7fe4c9dd3d7 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-09T15:03:26+08:00 | Bomeng Wang | wangbomeng@calculet.tech | 2026-03-09T15:03:26+08:00 | | clart: past_kv_len = n_tokens |
| 111 | c6ea3f0d2468e3b8eab5de80bacee7fe4c9dd3d7 | f8fba66b657c51a1df1efc7267bfea9f6140cebf | wangbomeng | wangbomeng@calculet.tech | 2026-03-09T03:26:47+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-09T03:26:47+08:00 | | CMakeLists: STATIC to SHARED |
| 112 | f8fba66b657c51a1df1efc7267bfea9f6140cebf | db4a61d2234c2f457b42ecc3f7219a1e1a35508a | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T15:06:36+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T07:05:39+08:00 | | fix calrt_install include path |
| 113 | db4a61d2234c2f457b42ecc3f7219a1e1a35508a | e672e02ca1a582f9e3a9d011d9e1b4fea4942a25 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T06:59:58+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T06:59:58+08:00 | | adapt to host-rt install structure |
| 114 | e672e02ca1a582f9e3a9d011d9e1b4fea4942a25 | 27b82f64252779ad160a3fa76bd378c39421de36 7a3a9a0e76bb5f530beafcfe52e87e388e93117a | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T11:47:09+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T11:47:09+08:00 | | Merge branch 'dev' of http://192.168.10.26:1000/yunzhe/llama.cpp into dev |
| 115 | 27b82f64252779ad160a3fa76bd378c39421de36 | 0bd95719fed2f0247c8d0199dd897d24157dfae6 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T11:42:45+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-03-06T11:45:03+08:00 | | adapt to calrt_objs target |
| 116 | 7a3a9a0e76bb5f530beafcfe52e87e388e93117a | 583c387efaf7bc030574e554088de79d09a27fbd | wangbomeng | wangbomeng@calculet.tech | 2026-03-05T07:24:59+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-03-05T07:24:59+08:00 | | calrt_infer: fix past_kv_len |
| 117 | 583c387efaf7bc030574e554088de79d09a27fbd | 0edc14a60a2d5596f71d728c655e3523da924d89 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T02:24:56+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T02:24:56+08:00 | | caserver: lim generated tokens to n_ctx_per_seq |
| 118 | 0edc14a60a2d5596f71d728c655e3523da924d89 | aa0a17d719326a63e0d6fea60aa0ced44e7abc33 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:57:39+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:57:39+08:00 | | caserver: limit generated tokens to n_ctx_per_seq |
| 119 | aa0a17d719326a63e0d6fea60aa0ced44e7abc33 | fa332e773da681f3e77d6ffe83a87edb557f758b | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:40:11+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:40:11+08:00 | | calrt: add bf16_to_fp32 |
| 120 | fa332e773da681f3e77d6ffe83a87edb557f758b | aa54a7c637f2c92d924a5f35386975b1c27de898 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:31:42+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:31:42+08:00 | | Add bf16_to_fp32 |
| 121 | aa54a7c637f2c92d924a5f35386975b1c27de898 | 84f2c0b5f334a7ab77347c6779b6027a951d0ff0 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:30:36+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-27T01:30:36+08:00 | | Add bf16_to_fp32 |
| 122 | 84f2c0b5f334a7ab77347c6779b6027a951d0ff0 | 0bd95719fed2f0247c8d0199dd897d24157dfae6 | wangbomeng | wangbomeng@calculet.tech | 2026-02-09T01:46:36+08:00 | wangbomeng | wangbomeng@calculet.tech | 2026-02-09T01:46:36+08:00 | | add bf16_to_fp32 in untile_decode/prefill_cpy |
| 123 | 0bd95719fed2f0247c8d0199dd897d24157dfae6 | 365eb0b719e30b0f9e348d854f11c9c3f954a96e | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-07T10:25:06+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-07T10:25:06+08:00 | | fix get_model_by_name |
| 124 | 365eb0b719e30b0f9e348d854f11c9c3f954a96e | 886cd1fb3b217efcf51c46209fcb694022346797 | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-07T09:05:17+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-07T09:05:17+08:00 | | comment out pos buffer file |
| 125 | 886cd1fb3b217efcf51c46209fcb694022346797 | 9328f21378275f2354a19c3af907332b216ab492 | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-06T17:36:54+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-02-06T17:36:54+08:00 | | fix compile problem |
| 126 | 9328f21378275f2354a19c3af907332b216ab492 | 49fea70a20f550bfdb59d0ffe041b54919c22447 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-28T15:32:36+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-28T15:32:36+08:00 | | add prefill-only mode |
| 127 | 49fea70a20f550bfdb59d0ffe041b54919c22447 | 74a118df31b6a4acc8876b2a2db90d2395cc6ec3 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-28T09:11:16+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-28T09:11:16+08:00 | | adapt to new calbin test_case |
| 128 | 74a118df31b6a4acc8876b2a2db90d2395cc6ec3 | 98e8f9acfe87cd93646482cb2d942b8a132f0852 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-13T14:32:47+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2026-01-13T14:32:47+08:00 | | sync with calrt update |
| 129 | 98e8f9acfe87cd93646482cb2d942b8a132f0852 | 8c8152a04036bb0d4756a4e08e54f514dae4e64d | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-10T00:20:57+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-10T00:20:57+08:00 | | add timestamp |
| 130 | 8c8152a04036bb0d4756a4e08e54f514dae4e64d | 5d2387931528fbe04aa8d66de935147efb65ca4d | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-09T14:40:04+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-09T14:40:28+08:00 | | adapt to calrt csr |
| 131 | 5d2387931528fbe04aa8d66de935147efb65ca4d | e6285a748657add8d974b78ca920c5b41753ee12 | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-08T16:40:34+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-12-08T16:40:54+08:00 | | add timestamp for one decode process |
| 132 | e6285a748657add8d974b78ca920c5b41753ee12 | 6b1ecdbbe3cc72ed8888be38f9e934cfd79b9fbb | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-27T16:58:35+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-27T16:58:35+08:00 | | adapt to calrt |
| 133 | 6b1ecdbbe3cc72ed8888be38f9e934cfd79b9fbb | 240d9bb75241c91896afe125f2fc74698a7395f3 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-25T17:06:14+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-25T17:06:14+08:00 | | add kv_tensor for pld test |
| 134 | 240d9bb75241c91896afe125f2fc74698a7395f3 | 7eeb241aad429a68fb8cfa3ba85d000fb2348cd2 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-24T12:03:00+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-24T12:03:00+08:00 | | adapt to calrt modification |
| 135 | 7eeb241aad429a68fb8cfa3ba85d000fb2348cd2 | dfb6250031a25058dbc3757e8a7ec75d90dcb48a | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-21T09:20:01+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-21T09:20:01+08:00 | | fix byte_size error |
| 136 | dfb6250031a25058dbc3757e8a7ec75d90dcb48a | e31bce25e85dc7fdf0a0e69fd0e58ba4a2ffaf3e | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-18T10:28:00+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-18T10:43:01+08:00 | | add elf in CMakelist |
| 137 | e31bce25e85dc7fdf0a0e69fd0e58ba4a2ffaf3e | 6444e227c36df16199c40d8cead2244fe014f377 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-07T11:52:17+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-07T11:52:17+08:00 | | dump src tensor |
| 138 | 6444e227c36df16199c40d8cead2244fe014f377 | d81d7b190ff5f21325167936496dfe5ab48b922e | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-07T09:19:11+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-07T09:19:11+08:00 | | test golden case |
| 139 | d81d7b190ff5f21325167936496dfe5ab48b922e | e492f5dc8cec17529f53ef2bbdf7b502107f8148 1f5accb8d0056e6099cd5b772b1cb787dd590a13 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-04T14:43:32+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-04T15:04:07+08:00 | | Merge branch 'master' into dev |
| 140 | e492f5dc8cec17529f53ef2bbdf7b502107f8148 | 2e05afd22b3d77c7013b11eee88e88b17188f21a | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-04T14:09:59+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-11-04T14:09:59+08:00 | | redo untile logic |
| 141 | 1f5accb8d0056e6099cd5b772b1cb787dd590a13 | 2759ccdb4adc8568add4316780d5e675519b0775 | Noah | 99681487+NoahOksuz@users.noreply.github.com | 2025-11-04T05:04:59Z | GitHub | noreply@github.com | 2025-11-03T21:04:59-08:00 | | Fix garbled output with REPACK at high thread counts (#16956) |
| 142 | 2759ccdb4adc8568add4316780d5e675519b0775 | c5023daf607c578d6344c628eb7da18ac3d92d32 | Aman Gupta | amangupta052@gmail.com | 2025-11-04T10:53:48+08:00 | GitHub | noreply@github.com | 2025-11-04T10:53:48+08:00 | | CUDA: avoid mul + bias fusion when doing fusion (#16935) |
| 143 | c5023daf607c578d6344c628eb7da18ac3d92d32 | e7da30b584dc1f2ee0414c4a1298ce64eef97e8d | lhez | lih@qti.qualcomm.com | 2025-11-03T11:47:57-08:00 | GitHub | noreply@github.com | 2025-11-03T11:47:57-08:00 | | opencl: support imrope (#16914) |
| 144 | e7da30b584dc1f2ee0414c4a1298ce64eef97e8d | ed8aa63320393512bdcfe4b05b5ae01ba91888e1 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-11-03T18:53:26+01:00 | GitHub | noreply@github.com | 2025-11-03T18:53:26+01:00 | | fix: Viewing multiple PDF attachments (#16974) |
| 145 | ed8aa63320393512bdcfe4b05b5ae01ba91888e1 | 48bd26501b08a3f0bff1249db47f313641f7bebb | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-11-03T18:01:59+01:00 | GitHub | noreply@github.com | 2025-11-03T18:01:59+01:00 | | model-conversion : pass config to from_pretrained (#16963) |
| 146 | 48bd26501b08a3f0bff1249db47f313641f7bebb | 622cd010ff4a65bef67edbf2f9bf4707c01f98f7 | Georgi Gerganov | ggerganov@gmail.com | 2025-11-03T15:38:23+02:00 | GitHub | noreply@github.com | 2025-11-03T14:38:23+01:00 | | server : add props.model_alias (#16943) |
| 147 | 622cd010ff4a65bef67edbf2f9bf4707c01f98f7 | 070ff4d5356083d60b807bb34d36b31c3653a29e | theo77186 | theo77186@users.noreply.github.com | 2025-11-03T14:29:11+01:00 | GitHub | noreply@github.com | 2025-11-03T14:29:11+01:00 | | ggml: CUDA: add head size 72 for flash-attn (#16962) |
| 148 | 070ff4d5356083d60b807bb34d36b31c3653a29e | bf7b0c9725ed0c406c5debaea022ec7258430634 | Xuan-Son Nguyen | son@huggingface.co | 2025-11-03T11:11:18+01:00 | GitHub | noreply@github.com | 2025-11-03T11:11:18+01:00 | | mtmd: add --image-min/max-tokens (#16921) |
| 149 | bf7b0c9725ed0c406c5debaea022ec7258430634 | fcfce040e816c542f374ad51461b3561d73c4bc9 | Xuan-Son Nguyen | son@huggingface.co | 2025-11-03T10:25:55+01:00 | GitHub | noreply@github.com | 2025-11-03T10:25:55+01:00 | | mtmd: pad mask for qwen2.5vl (#16954) |
| 150 | fcfce040e816c542f374ad51461b3561d73c4bc9 | ee3a5a10adf9e83722d1914dddc56a0623ececaf | Jinyang He | hejinyang@loongson.cn | 2025-11-03T14:40:02+08:00 | GitHub | noreply@github.com | 2025-11-03T08:40:02+02:00 | | ggml : LoongArch fixes (#16958) |
| 151 | ee3a5a10adf9e83722d1914dddc56a0623ececaf | 7e994168b1ccc12337ba8de939c4fd466107c1fb | Olivier Chafik | olivier.chafik@gmail.com | 2025-11-03T05:33:56Z | GitHub | noreply@github.com | 2025-11-03T07:33:56+02:00 | | sync: minja (glm 4.6 & minmax m2 templates) (#16949) |
| 152 | 7e994168b1ccc12337ba8de939c4fd466107c1fb | bcfa87622ae46be6345a8e3dfdbdc5ba5414042b | shani-f | s0556787439@gmail.com | 2025-11-03T03:35:33+02:00 | GitHub | noreply@github.com | 2025-11-03T09:35:33+08:00 | | SYCL: optimized repeat_back kernel (3× fewer asm instructions, 2× faster)Feature/sycl repeat back opt (#16869) |
| 153 | bcfa87622ae46be6345a8e3dfdbdc5ba5414042b | a2054e3a8ff0da3978a4acc18c349ff58554d336 | Sascha Rogmann | 59577610+srogmann@users.noreply.github.com | 2025-11-03T00:41:08+01:00 | GitHub | noreply@github.com | 2025-11-03T00:41:08+01:00 | | feat(webui): improve LaTeX rendering with currency detection (#16508) |
| 154 | a2054e3a8ff0da3978a4acc18c349ff58554d336 | dd5286805004db1f9ac3176a1cbbfe373bdda0f8 | Shagun Bera | 141054835+notV3NOM@users.noreply.github.com | 2025-11-03T04:40:30+05:30 | GitHub | noreply@github.com | 2025-11-03T00:10:30+01:00 | | test-backend-ops : fix segfault in moe-expert-reduce test in support mode and coverage (#16936) |
| 155 | dd5286805004db1f9ac3176a1cbbfe373bdda0f8 | 6b9a52422bac0f50dd8f1f8386744fa3ce9783bf | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-11-02T23:11:21+01:00 | GitHub | noreply@github.com | 2025-11-02T23:11:21+01:00 | | ci : disable failing riscv cross build (#16952) |
| 156 | 6b9a52422bac0f50dd8f1f8386744fa3ce9783bf | 2f966b8ed87514e74bb96592217226cb6a6974dd | Zhiyong Wang | 85110830+ravenouse@users.noreply.github.com | 2025-11-02T13:08:04-08:00 | GitHub | noreply@github.com | 2025-11-02T22:08:04+01:00 | | model: add Janus Pro for image understanding (#16906) |
| 157 | 2f966b8ed87514e74bb96592217226cb6a6974dd | cd5e3b57541ecc52421130742f4d89acbcf77cd4 | Georgi Gerganov | ggerganov@gmail.com | 2025-11-02T22:21:48+02:00 | GitHub | noreply@github.com | 2025-11-02T21:21:48+01:00 | | clip : use FA (#16837) |
| 158 | cd5e3b57541ecc52421130742f4d89acbcf77cd4 | 87c9efc3b297b8a498716b1db3d061842e6fc85b | Georgi Gerganov | ggerganov@gmail.com | 2025-11-02T18:14:04+02:00 | GitHub | noreply@github.com | 2025-11-02T18:14:04+02:00 | | server : support unified cache across slots (#16736) |
| 159 | 87c9efc3b297b8a498716b1db3d061842e6fc85b | 76af40aaaad78c42faecd8016a88362c788b84b0 | Aldehir Rojas | hello@alde.dev | 2025-11-02T08:56:28-06:00 | GitHub | noreply@github.com | 2025-11-02T16:56:28+02:00 | | common : move gpt-oss reasoning processing to init params (#16937) |
| 160 | 76af40aaaad78c42faecd8016a88362c788b84b0 | 7db35a7958a943be1693879f42d166f152613979 | Adrian Lundberg | 47256989+alundb@users.noreply.github.com | 2025-11-02T10:28:37+01:00 | GitHub | noreply@github.com | 2025-11-02T11:28:37+02:00 | | docs: remove llama_sampler_accept reference in sampling sample usage (#16920) |
| 161 | 7db35a7958a943be1693879f42d166f152613979 | a864132ba546b6385922226480ee4e392aaa065c | mnehete32 | 33429707+mnehete32@users.noreply.github.com | 2025-11-02T08:42:57+05:30 | GitHub | noreply@github.com | 2025-11-02T11:12:57+08:00 | | CUDA: add FLOOR, CEIL, ROUND, TRUNC unary ops (#16917) |
| 162 | a864132ba546b6385922226480ee4e392aaa065c | d38d9f0877a5872daa3c5f06fb9a86376bf15d50 | Aaron Teo | aaron.teo1@ibm.com | 2025-11-02T08:48:46+08:00 | GitHub | noreply@github.com | 2025-11-02T08:48:46+08:00 | | devops: fix failing s390x docker build (#16918) |
| 163 | d38d9f0877a5872daa3c5f06fb9a86376bf15d50 | 7fd205a8e8832ead273f62c5fd81d2b8aa910535 | Aaron Teo | aaron.teo1@ibm.com | 2025-11-02T08:48:23+08:00 | GitHub | noreply@github.com | 2025-11-02T08:48:23+08:00 | | ggml: add s390x cpu-feats (#16774) |
| 164 | 7fd205a8e8832ead273f62c5fd81d2b8aa910535 | 2f68ce7cfd20e9e7098514bf730e5389b7bba908 | Georgi Gerganov | ggerganov@gmail.com | 2025-11-02T00:15:31+02:00 | GitHub | noreply@github.com | 2025-11-02T00:15:31+02:00 | | scripts : add script to bench models (#16894) |
| 165 | 2f68ce7cfd20e9e7098514bf730e5389b7bba908 | e4a71599e5846110159955dec0008eb4aa24222b | Pascal | admin@serveurperso.com | 2025-11-01T19:49:51+01:00 | GitHub | noreply@github.com | 2025-11-01T19:49:51+01:00 | | webui: auto-refresh /props on inference start to resync model metadata (#16784) |
| 166 | e4a71599e5846110159955dec0008eb4aa24222b | dd5e8cab512c7752392b6e51a6f118f348fb3f16 | Pascal | admin@serveurperso.com | 2025-11-01T17:14:54+01:00 | GitHub | noreply@github.com | 2025-11-01T17:14:54+01:00 | | webui: add HTML/JS preview support to MarkdownContent with sandboxed iframe (#16757) |
| 167 | dd5e8cab512c7752392b6e51a6f118f348fb3f16 | cf659bbb8ef9eb048e5153a27cf787fd83c05560 | Adrien Gallouët | angt@huggingface.co | 2025-11-01T16:52:17+01:00 | GitHub | noreply@github.com | 2025-11-01T16:52:17+01:00 | | vendor : update cpp-httplib to 0.27.0 (#16846) |
| 168 | cf659bbb8ef9eb048e5153a27cf787fd83c05560 | d8b860a219c2415faac8cc0e50b48b4aa11e3b64 | Xuan-Son Nguyen | son@huggingface.co | 2025-11-01T15:51:36+01:00 | GitHub | noreply@github.com | 2025-11-01T15:51:36+01:00 | | mtmd: refactor preprocessing + support max/min pixels (#16878) |
| 169 | d8b860a219c2415faac8cc0e50b48b4aa11e3b64 | 1ae74882f8f6755e44dff8f23f3abdc5b53ab7c1 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-11-01T15:35:57+01:00 | GitHub | noreply@github.com | 2025-11-01T15:35:57+01:00 | | Add a setting to display message generation statistics (#16901) |
| 170 | 1ae74882f8f6755e44dff8f23f3abdc5b53ab7c1 | 961660b8c395afb8903f61ace69396a158ea0cb4 | Jaromír Hradílek | jhradilek@gmail.com | 2025-11-01T15:02:57+01:00 | GitHub | noreply@github.com | 2025-11-01T15:02:57+01:00 | | webui: recognize AsciiDoc files as valid text files (#16850) |
| 171 | 961660b8c395afb8903f61ace69396a158ea0cb4 | 74fef4129fa58f6277c243777a4894371479dbad | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-11-01T11:01:42+01:00 | GitHub | noreply@github.com | 2025-11-01T11:01:42+01:00 | | common : allow --system-prompt-file for diffusion-cli (#16903) |
| 172 | 74fef4129fa58f6277c243777a4894371479dbad | 5d8bb900bc7daa84bfa7bb1d25ab7e32394919f3 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-11-01T08:55:25+01:00 | GitHub | noreply@github.com | 2025-11-01T09:55:25+02:00 | | codeowners : update after refactor (#16905) |
| 173 | 5d8bb900bc7daa84bfa7bb1d25ab7e32394919f3 | 2e76e013600cb0d51ccf158571ca1d0502952a07 | Jeff Bolz | jbolz@nvidia.com | 2025-11-01T00:52:14-05:00 | GitHub | noreply@github.com | 2025-11-01T06:52:14+01:00 | | vulkan: Fix multi_add invalid descriptor usage (#16899) |
| 174 | 2e76e013600cb0d51ccf158571ca1d0502952a07 | d3dc9dd898be805c23a408cc36daed5b3bf29221 | Jeff Bolz | jbolz@nvidia.com | 2025-11-01T00:45:28-05:00 | GitHub | noreply@github.com | 2025-11-01T06:45:28+01:00 | | vulkan: fuse mul_mat+add and mul_mat_id+add_id (#16868) |
| 175 | d3dc9dd898be805c23a408cc36daed5b3bf29221 | bea04522ff1a0d8559ccfd353aa018dcfbb608cc | Oliver Simons | osimons@nvidia.com | 2025-11-01T06:13:26+01:00 | GitHub | noreply@github.com | 2025-11-01T13:13:26+08:00 | | CUDA: Remove unneded bias/gate dims in fused mmvq (#16858) |
| 176 | bea04522ff1a0d8559ccfd353aa018dcfbb608cc | 0de0a01576772032008a689afc4d7c80685074c4 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-10-31T23:40:23+01:00 | GitHub | noreply@github.com | 2025-10-31T23:40:23+01:00 | | refactor : llama-model.cpp (#16252) |
| 177 | 0de0a01576772032008a689afc4d7c80685074c4 | e58d585604bbbbbc24f8149effe444621b5191c6 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-10-31T21:20:47+01:00 | GitHub | noreply@github.com | 2025-10-31T21:20:47+01:00 | | model : Minimax M2 (#16831) |
| 178 | e58d585604bbbbbc24f8149effe444621b5191c6 | 31c511a968348281e11d590446bb815048a1e912 | Giuseppe Scrivano | gscrivan@redhat.com | 2025-10-31T21:20:07+01:00 | GitHub | noreply@github.com | 2025-10-31T21:20:07+01:00 | | model : add Granite Hybrid nano types (#16896) |
| 179 | 31c511a968348281e11d590446bb815048a1e912 | 6d39015a744b47bb7643ea73f412e6d64677f305 | Johannes Gäßler | johannesg@5d6.de | 2025-10-31T15:57:19+01:00 | GitHub | noreply@github.com | 2025-10-31T15:57:19+01:00 | | CUDA: Volta tensor core support for MMF (#16843) |
| 180 | 6d39015a744b47bb7643ea73f412e6d64677f305 | 4146d6a1a6228711a487a1e3e9ddd120f8d027d7 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-31T16:25:50+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-31T16:26:28+02:00 | | sync : ggml |
| 181 | 4146d6a1a6228711a487a1e3e9ddd120f8d027d7 | 8da3c0e200a586f768ada6f38745acb01380174c | Aman Gupta | amangupta052@gmail.com | 2025-10-31T20:05:07+08:00 | GitHub | noreply@github.com | 2025-10-31T20:05:07+08:00 | | CUDA: add expert reduce kernel (#16857) |
| 182 | 8da3c0e200a586f768ada6f38745acb01380174c | c22473b580807929fd9e3a3344a48e8cfbe6c88f | Georgi Gerganov | ggerganov@gmail.com | 2025-10-31T13:50:33+02:00 | GitHub | noreply@github.com | 2025-10-31T13:50:33+02:00 | | batch : fix consistency checks for the input positions (#16890) |
| 183 | c22473b580807929fd9e3a3344a48e8cfbe6c88f | 0f715b4e759acceccb9f437cfd2a988fff85514a | Georgi Gerganov | ggerganov@gmail.com | 2025-10-31T10:54:19+02:00 | GitHub | noreply@github.com | 2025-10-31T10:54:19+02:00 | | server : don't print user inputs to console (#16871) |
| 184 | 0f715b4e759acceccb9f437cfd2a988fff85514a | d2d931f173b8a736b08999436e9259aafddec718 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-31T09:51:26+01:00 | GitHub | noreply@github.com | 2025-10-31T09:51:26+01:00 | | server : fix typos in server.cpp comments [no ci] (#16883) |
| 185 | d2d931f173b8a736b08999436e9259aafddec718 | 2976b0374d36609b0429dd6ce48807e2ad39a7c2 | Jeff Bolz | jbolz@nvidia.com | 2025-10-31T02:34:47-05:00 | GitHub | noreply@github.com | 2025-10-31T08:34:47+01:00 | | vulkan: disable spirv-opt for rope shaders (#16872) |
| 186 | 2976b0374d36609b0429dd6ce48807e2ad39a7c2 | d2a2673dd1c86a01ab010e18b13b8bb959968c48 | Masato Nakasaka | masato.nakasaka@intel.com | 2025-10-31T16:18:59+09:00 | GitHub | noreply@github.com | 2025-10-31T08:18:59+01:00 | | vulkan: Fix crash when FP16 mul_mat accumulation is not supported (#16796) |
| 187 | d2a2673dd1c86a01ab010e18b13b8bb959968c48 | 13002a08960e51a76c4d696165b5d7638d2f2b99 | Ruben Ortlam | picard12@live.de | 2025-10-31T08:14:49+01:00 | GitHub | noreply@github.com | 2025-10-31T08:14:49+01:00 | | vulkan: fix shmem overrun in mmq id shader (#16873) |
| 188 | 2e05afd22b3d77c7013b11eee88e88b17188f21a | c20c0a5c0c44179f5267a262546624854b1203d1 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-31T14:06:06+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-31T14:06:06+08:00 | | add pcie support |
| 189 | 13002a08960e51a76c4d696165b5d7638d2f2b99 | 6eb208d17ea29bb60295d9a2b5e7122dfb8f4b55 | l3utterfly | gc.pthzfoldr@gmail.com | 2025-10-31T12:46:31+08:00 | GitHub | noreply@github.com | 2025-10-30T21:46:31-07:00 | | ggml-hexagon: respect input size when getting/setting tensor data (#16836) |
| 190 | 6eb208d17ea29bb60295d9a2b5e7122dfb8f4b55 | 9984cbb61d12491d604484fb38b678fa15064061 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-31T00:34:27+01:00 | GitHub | noreply@github.com | 2025-10-31T00:34:27+01:00 | | ci : enable free-disk-space on cuda docker build (#16877) |
| 191 | 9984cbb61d12491d604484fb38b678fa15064061 | ce18efeaf1234bd08f25bc08f88ef95cf5d9e51f | lhez | lih@qti.qualcomm.com | 2025-10-30T16:00:20-07:00 | GitHub | noreply@github.com | 2025-10-30T16:00:20-07:00 | | opencl: fix boundary handling for mul_mm (#16875) |
| 192 | ce18efeaf1234bd08f25bc08f88ef95cf5d9e51f | 16724b5b6836a2d4b8936a5824d2ff27c52b4517 | RodriMora | bullerwins@gmail.com | 2025-10-30T23:15:03+01:00 | GitHub | noreply@github.com | 2025-10-30T23:15:03+01:00 | | convert : update transformers requirements (#16866) |
| 193 | 16724b5b6836a2d4b8936a5824d2ff27c52b4517 | b52edd25586fabb70f0c21b274473b307cf14499 | chansikpark | chansik.park@gmail.com | 2025-10-30T14:22:23-04:00 | GitHub | noreply@github.com | 2025-10-30T20:22:23+02:00 | | server : bump request URI max length to 32768 (#16862) |
| 194 | b52edd25586fabb70f0c21b274473b307cf14499 | 517b7170e1a4d733583c4b07c5b7a49acc05911c | Georgi Gerganov | ggerganov@gmail.com | 2025-10-30T18:42:57+02:00 | GitHub | noreply@github.com | 2025-10-30T18:42:57+02:00 | | server : remove n_past (#16818) |
| 195 | 517b7170e1a4d733583c4b07c5b7a49acc05911c | 835e918d8428f5119927d7150bf5a26176dedda0 | Max Krasnyansky | maxk@qti.qualcomm.com | 2025-10-30T09:06:13-07:00 | GitHub | noreply@github.com | 2025-10-30T09:06:13-07:00 | | cpu: introduce chunking for repack matmuls and enable matmul-id chunking on ARM64 (#16833) |
| 196 | 835e918d8428f5119927d7150bf5a26176dedda0 | d261223d24e97f2df50220e4a5b7f0adb69bba81 | Shagun Bera | 141054835+notV3NOM@users.noreply.github.com | 2025-10-30T21:17:31+05:30 | GitHub | noreply@github.com | 2025-10-30T17:47:31+02:00 | | common: fix typo in cli help text (#16864) |
| 197 | d261223d24e97f2df50220e4a5b7f0adb69bba81 | dcca0d3ab840ebe9b2ccd4719033d408eeb758d7 | JJJYmmm | 92386084+JJJYmmm@users.noreply.github.com | 2025-10-30T23:19:14+08:00 | GitHub | noreply@github.com | 2025-10-30T16:19:14+01:00 | | model: add support for qwen3vl series (#16780) |
| 198 | dcca0d3ab840ebe9b2ccd4719033d408eeb758d7 | bacddc049a00786df44e682262f6e298742bfbc3 | Max Krasnyansky | maxk@qti.qualcomm.com | 2025-10-30T05:26:05-07:00 | GitHub | noreply@github.com | 2025-10-30T14:26:05+02:00 | | cpu: introduce chunking for flash attention (#16829) |
| 199 | bacddc049a00786df44e682262f6e298742bfbc3 | 229bf686287d18f82c44e89888cc662145ecfdb4 | Tianyue-Zhao | zhaotianyue@outlook.com | 2025-10-30T07:18:50-04:00 | GitHub | noreply@github.com | 2025-10-30T12:18:50+01:00 | | model: Add support for CogVLM model (#15002) |
| 200 | 229bf686287d18f82c44e89888cc662145ecfdb4 | d7395115baf395b75a73a17b0b796e746e468da9 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-30T08:56:28+01:00 | GitHub | noreply@github.com | 2025-10-30T08:56:28+01:00 | | cuda : fix argsort with 64k+ rows (#16849) |
| 201 | d7395115baf395b75a73a17b0b796e746e468da9 | 052df28b0e3ca0398e5928c61c7de40254317894 | Jan Boon | jan.boon@kaetemi.be | 2025-10-30T14:30:58+08:00 | GitHub | noreply@github.com | 2025-10-30T08:30:58+02:00 | | llama : use std::abs instead of abs (#16853) |
| 202 | 052df28b0e3ca0398e5928c61c7de40254317894 | 8b11deea4663f29d3e042ce1056ba643264cd5f1 | Jeff Bolz | jbolz@nvidia.com | 2025-10-30T01:27:41-05:00 | GitHub | noreply@github.com | 2025-10-30T07:27:41+01:00 | | vulkan: Handle argsort with a large number of rows (#16851) |
| 203 | 8b11deea4663f29d3e042ce1056ba643264cd5f1 | b9ce94017729465895402cbcfffb51fa926c15e3 | Oliver Simons | osimons@nvidia.com | 2025-10-30T04:34:15+01:00 | GitHub | noreply@github.com | 2025-10-30T11:34:15+08:00 | | Hide latency of bias and gate-loading (#16847) |
| 204 | b9ce94017729465895402cbcfffb51fa926c15e3 | 3464bdac37027c5e9661621fc75ffcef3c19c6ef | Jeff Bolz | jbolz@nvidia.com | 2025-10-29T15:13:10-05:00 | GitHub | noreply@github.com | 2025-10-29T15:13:10-05:00 | | vulkan: Fuse rope+set_rows (#16769) |
| 205 | 3464bdac37027c5e9661621fc75ffcef3c19c6ef | e3af5563bd049141e036b50f843196db33d23e97 | Xuan-Son Nguyen | son@huggingface.co | 2025-10-29T20:11:39+01:00 | GitHub | noreply@github.com | 2025-10-29T20:11:39+01:00 | | llama: fix ASAN error with M-RoPE (#16848) |
| 206 | e3af5563bd049141e036b50f843196db33d23e97 | 10fcc41290e233788f5a4215314156e8e023eb92 | Xuan-Son Nguyen | son@huggingface.co | 2025-10-29T18:09:18+01:00 | GitHub | noreply@github.com | 2025-10-29T18:09:18+01:00 | | llama: store mrope data in KV cell (#16825) |
| 207 | 10fcc41290e233788f5a4215314156e8e023eb92 | bcf5bda6f5df559565d11d7c8e8295c1159a85ec | Jeff Bolz | jbolz@nvidia.com | 2025-10-29T08:44:29-05:00 | GitHub | noreply@github.com | 2025-10-29T14:44:29+01:00 | | vulkan: Update topk_moe fusion to handle gpt's late softmax (#16656) |
| 208 | bcf5bda6f5df559565d11d7c8e8295c1159a85ec | 3eb2be1ca5f37480aeb16102970d9e65f43347fe | Ruben Ortlam | picard12@live.de | 2025-10-29T14:39:03+01:00 | GitHub | noreply@github.com | 2025-10-29T14:39:03+01:00 | | Vulkan MMQ Integer Dot Refactor and K-Quant support (#16536) |
| 209 | 3eb2be1ca5f37480aeb16102970d9e65f43347fe | e41bcce8f0b53032a1fed275cd253e931c041cf6 | Max Krasnyansky | maxk@qti.qualcomm.com | 2025-10-29T06:29:12-07:00 | GitHub | noreply@github.com | 2025-10-29T06:29:12-07:00 | | Hexagon Op queue & dispatch optimizations (#16820) |
| 210 | e41bcce8f0b53032a1fed275cd253e931c041cf6 | 144a4ce824b6bd0e48d62009d10cae1daf5308db | Aman Gupta | amangupta052@gmail.com | 2025-10-29T21:11:53+08:00 | GitHub | noreply@github.com | 2025-10-29T21:11:53+08:00 | | CUDA: use fastdiv in set-rows (#16834) |
| 211 | 144a4ce824b6bd0e48d62009d10cae1daf5308db | f549b0007dbdd683215820f7229ce180a12b191d | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-29T14:09:50+01:00 | GitHub | noreply@github.com | 2025-10-29T14:09:50+01:00 | | vendor : sync minja (#16500) |
| 212 | f549b0007dbdd683215820f7229ce180a12b191d | 9a3ea685b937c0f0cbfda2e50004ea54bf187512 | Jeff Bolz | jbolz@nvidia.com | 2025-10-29T03:53:04-05:00 | GitHub | noreply@github.com | 2025-10-29T09:53:04+01:00 | | vulkan: Call ggml_vk_buffer_write_2d from ggml_vk_buffer_copy (#16793) |
| 213 | 9a3ea685b937c0f0cbfda2e50004ea54bf187512 | 338074c383c81366320d176d83b94b0a567ee0c2 | Aman Gupta | amangupta052@gmail.com | 2025-10-29T15:55:06+08:00 | GitHub | noreply@github.com | 2025-10-29T15:55:06+08:00 | | CUDA: Fix bug in topk-moe for gpt-oss (#16821) |
| 214 | 338074c383c81366320d176d83b94b0a567ee0c2 | 851553ea6b24cb39fd5fd188b437d777cb411de8 | YaelLogic | y0548591250@gmail.com | 2025-10-29T08:14:39+02:00 | GitHub | noreply@github.com | 2025-10-29T14:14:39+08:00 | | sycl: add RMS_NORM_BACK operation support (#16808) |
| 215 | 851553ea6b24cb39fd5fd188b437d777cb411de8 | 85a7d8677bf2200981e52f744a21d5267964ffcf | YaelGitAccount | 38328157276@mby.co.il | 2025-10-28T21:10:28+02:00 | GitHub | noreply@github.com | 2025-10-28T20:10:28+01:00 | | cuda: add SET operation support (#16804) |
| 216 | 85a7d8677bf2200981e52f744a21d5267964ffcf | a8ca18b4b815a2abdbecb958ee5f4c542d69aac7 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-28T20:19:44+02:00 | GitHub | noreply@github.com | 2025-10-28T20:19:44+02:00 | | memory : remove KV cache size padding (#16812) |
| 217 | a8ca18b4b815a2abdbecb958ee5f4c542d69aac7 | 8284efc35c909217ee1b9a06903245d808ac2283 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-28T19:41:43+02:00 | GitHub | noreply@github.com | 2025-10-28T19:41:43+02:00 | | llama-bench : clarify benchmarked parts of the computation (#16823) |
| 218 | 8284efc35c909217ee1b9a06903245d808ac2283 | 1c1409e13169d0ced4b1b4d39fb5b268b7525091 | l3utterfly | gc.pthzfoldr@gmail.com | 2025-10-28T23:16:20+08:00 | GitHub | noreply@github.com | 2025-10-28T08:16:20-07:00 | | initialise buffer.device in ggml_hexagon_session (#16816) |
| 219 | 1c1409e13169d0ced4b1b4d39fb5b268b7525091 | 7a0e900e3615fa46c074a7fdf900b47d3c0a1c7e | Sam Malayek | 12037535+SamMalayek@users.noreply.github.com | 2025-10-28T03:51:41-07:00 | GitHub | noreply@github.com | 2025-10-28T12:51:41+02:00 | | embedding: add raw option for --embd-output-format (#16541) |
| 220 | 7a0e900e3615fa46c074a7fdf900b47d3c0a1c7e | 280d97be9660e7a5feaa28a6e7a299bc73dd83fc | Johannes Gäßler | johannesg@5d6.de | 2025-10-28T11:23:54+01:00 | GitHub | noreply@github.com | 2025-10-28T11:23:54+01:00 | | llama: consistent ctx <-> buf order for KV cache (#16746) |
| 221 | 280d97be9660e7a5feaa28a6e7a299bc73dd83fc | 3479efd112b3910d2d008ad08fe4cb4895605903 | Aldehir Rojas | hello@alde.dev | 2025-10-28T03:37:52-05:00 | GitHub | noreply@github.com | 2025-10-28T09:37:52+01:00 | | grammar : support array references in json schema (#16792) |
| 222 | 3479efd112b3910d2d008ad08fe4cb4895605903 | 463bbf20bfbe0da10ac58984d4d62bed5a49362a | Chenguang Li | 757486878@qq.com | 2025-10-28T10:54:53+08:00 | GitHub | noreply@github.com | 2025-10-28T10:54:53+08:00 | | CANN: Improve device ID handling and aclnnArange checks (#16752) |
| 223 | 463bbf20bfbe0da10ac58984d4d62bed5a49362a | ad8d36beffd791db10c94eb9e964afb891e3ca55 | Aman Gupta | amangupta052@gmail.com | 2025-10-28T10:31:21+08:00 | GitHub | noreply@github.com | 2025-10-28T10:31:21+08:00 | | CUDA: add unused vars to mmvf and mmvq (#16807) |
| 224 | ad8d36beffd791db10c94eb9e964afb891e3ca55 | c053e18a66dd95dc340aa61317877c2a41d4e3cf | tamarPal | tamarp3385@gmail.com | 2025-10-28T03:50:33+02:00 | GitHub | noreply@github.com | 2025-10-28T09:50:33+08:00 | | sycl: add SSM_CONV operation support (#16800) |
| 225 | c053e18a66dd95dc340aa61317877c2a41d4e3cf | e1ab0848037c9f9bfe68b3e1cee9ee375e1018a3 | Yuri Khrustalev | ykhrustalev@users.noreply.github.com | 2025-10-27T18:54:01-04:00 | GitHub | noreply@github.com | 2025-10-27T23:54:01+01:00 | | chat: Add LFM2 tool handling (#16763) |
| 226 | e1ab0848037c9f9bfe68b3e1cee9ee375e1018a3 | 5a4ff43e7dd049e35942bc3d12361dab2f155544 | Xuan-Son Nguyen | son@huggingface.co | 2025-10-27T23:12:16+01:00 | GitHub | noreply@github.com | 2025-10-27T23:12:16+01:00 | | mtmd : fix idefics3 preprocessing (#16806) |
| 227 | 5a4ff43e7dd049e35942bc3d12361dab2f155544 | 10640e31aab0819f31c1e1f2d008b019ee737232 | Diego Devesa | slarengh@gmail.com | 2025-10-27T13:51:28-07:00 | GitHub | noreply@github.com | 2025-10-27T21:51:28+01:00 | | llama : disable pipeline parallelism if compute buffer allocation fails (#16748) |
| 228 | 10640e31aab0819f31c1e1f2d008b019ee737232 | 80d28f104c0c3e61c11d6af073642b246a9fc19c | Acly | aclysia@gmail.com | 2025-10-27T21:50:22+01:00 | GitHub | noreply@github.com | 2025-10-27T21:50:22+01:00 | | ggml : fix interpolate with align-corners and ne=1 (#16700) |
| 229 | 80d28f104c0c3e61c11d6af073642b246a9fc19c | c55d53acec864f64afa1ba92972203dce1bf88f5 | Johannes Gäßler | johannesg@5d6.de | 2025-10-27T21:39:49+01:00 | GitHub | noreply@github.com | 2025-10-27T21:39:49+01:00 | | HIP: fix AMDGPU_TARGETS, update documentation (#16803) |
| 230 | c55d53acec864f64afa1ba92972203dce1bf88f5 | 945501f5ea4b8ca56b181ecb035e9ee3fb31f432 | Xuan-Son Nguyen | son@huggingface.co | 2025-10-27T16:02:58+01:00 | GitHub | noreply@github.com | 2025-10-27T16:02:58+01:00 | | model : add LightOnOCR-1B model (#16764) |
| 231 | 945501f5ea4b8ca56b181ecb035e9ee3fb31f432 | 75cbdd3fce38ea12d50cd19e73a069aa5dbbd5fa | Johannes Gäßler | johannesg@5d6.de | 2025-10-27T09:17:31+01:00 | GitHub | noreply@github.com | 2025-10-27T09:17:31+01:00 | | llama: fix leaked buffers for mmap + split files (#16765) |
| 232 | 75cbdd3fce38ea12d50cd19e73a069aa5dbbd5fa | 2b9bd9bf4e759c05db629ec1c391dc8aeaa71887 | Aman Gupta | amangupta052@gmail.com | 2025-10-27T09:25:10+08:00 | GitHub | noreply@github.com | 2025-10-27T09:25:10+08:00 | | test-backend-ops: print failed tests at the end (#16785) |
| 233 | 2b9bd9bf4e759c05db629ec1c391dc8aeaa71887 | 59fc1ec8e83b14354c1a3a8acf8c5c2cbf9af42f | tamarPal | tamarp3385@gmail.com | 2025-10-27T03:20:24+02:00 | GitHub | noreply@github.com | 2025-10-27T09:20:24+08:00 | | sycl: add ROLL operation support (#16665) |
| 234 | 59fc1ec8e83b14354c1a3a8acf8c5c2cbf9af42f | 75d33b9302f84a5b89f82205d2bcd8def5a64e0a | shani-f | s0556787439@gmail.com | 2025-10-27T03:19:50+02:00 | GitHub | noreply@github.com | 2025-10-27T09:19:50+08:00 | | sycl: add REPEAT_BACK operation support (#16734) |
| 235 | 75d33b9302f84a5b89f82205d2bcd8def5a64e0a | 3470a5c891dcc94363e492a3760af92b6b07241c | Aman Gupta | amangupta052@gmail.com | 2025-10-27T09:06:16+08:00 | GitHub | noreply@github.com | 2025-10-27T09:06:16+08:00 | | CUDA: support for weight clamp in top-k norm (#16702) |
| 236 | 3470a5c891dcc94363e492a3760af92b6b07241c | bd562fe4f7bd55625511d5f9d639c4fb1db1d440 | Acly | aclysia@gmail.com | 2025-10-26T23:19:03+01:00 | GitHub | noreply@github.com | 2025-10-26T23:19:03+01:00 | | ggml-alloc : make gallocr prefer chunks that allow memory reuse (#16788) |
| 237 | bd562fe4f7bd55625511d5f9d639c4fb1db1d440 | bbac6a26b2bd7f7c1f0831cb1e7b52734c66673b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-26T21:31:41+01:00 | GitHub | noreply@github.com | 2025-10-26T21:31:41+01:00 | | cuda : use fast copy when src and dst are of different type and contiguous (#16789) |
| 238 | bbac6a26b2bd7f7c1f0831cb1e7b52734c66673b | 73a48c9790d320476b3e5ef75bda09f2f8269e6e | leejet | leejet714@gmail.com | 2025-10-27T02:13:31+08:00 | GitHub | noreply@github.com | 2025-10-26T19:13:31+01:00 | | ggml: fix cuda kernel launch configuration for k_compute_batched_ptrs to support large batch (#16744) |
| 239 | 73a48c9790d320476b3e5ef75bda09f2f8269e6e | f696428ce8e4d16c17acbffeaa7feac3b0fb9061 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-26T17:21:23+01:00 | GitHub | noreply@github.com | 2025-10-26T17:21:23+01:00 | | convert : enable expert group selection for all models with it (#16691) |
| 240 | f696428ce8e4d16c17acbffeaa7feac3b0fb9061 | 7cce4f8158f0c4c88d8dadd4c23d33938127b897 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-26T17:20:32+01:00 | GitHub | noreply@github.com | 2025-10-26T17:20:32+01:00 | | graph : add clamping to ffn_moe_weights_sum to avoid div-by-zero (#16655) |
| 241 | 7cce4f8158f0c4c88d8dadd4c23d33938127b897 | 8d8862829cd770fc9c0c7a726ed162de824ff5ea | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-26T16:08:52+01:00 | GitHub | noreply@github.com | 2025-10-26T16:08:52+01:00 | | model : set res->t_embd in SmallThinker models (#16782) |
| 242 | 8d8862829cd770fc9c0c7a726ed162de824ff5ea | f77c13b91f4d25754b6a0b857f98a6bc922a0aa7 | amirai21 | 89905406+amirai21@users.noreply.github.com | 2025-10-26T14:01:20+02:00 | GitHub | noreply@github.com | 2025-10-26T13:01:20+01:00 | | docs : add Jamba to Text-only models list (#16778) |
| 243 | f77c13b91f4d25754b6a0b857f98a6bc922a0aa7 | 3cfa9c3f125763305b4226bc032f1954f08990dc | Aman Gupta | amangupta052@gmail.com | 2025-10-26T19:28:04+08:00 | GitHub | noreply@github.com | 2025-10-26T19:28:04+08:00 | | CUDA: General GEMV fusion (#16715) |
| 244 | 3cfa9c3f125763305b4226bc032f1954f08990dc | 5d195f17bc60eacc15cfb929f9403cf29ccdf419 | Gilad S. | 7817232+giladgd@users.noreply.github.com | 2025-10-26T06:37:38+02:00 | GitHub | noreply@github.com | 2025-10-26T05:37:38+01:00 | | vulkan: deduplicate Microsoft Direct3D12 devices (#16689) |
| 245 | 5d195f17bc60eacc15cfb929f9403cf29ccdf419 | 226f295f4dd92ad714533adc5497afed5fa88bb8 | Galunid | karolek1231456@gmail.com | 2025-10-25T20:41:36+02:00 | GitHub | noreply@github.com | 2025-10-25T20:41:36+02:00 | | convert : handle mmproj filename/path properly (#16760) |
| 246 | 226f295f4dd92ad714533adc5497afed5fa88bb8 | f90b4a8efe4466215c51155a5c22c1ae6207da23 | Shunta Saito | shunta.saito@gmail.com | 2025-10-25T19:26:27+09:00 | GitHub | noreply@github.com | 2025-10-25T12:26:27+02:00 | | model : set res->t_embd in PLaMo2 models (#16766) |
| 247 | f90b4a8efe4466215c51155a5c22c1ae6207da23 | 8423d019318b446640bb620e4ce80066d8530f05 | Giuseppe Scrivano | gscrivan@redhat.com | 2025-10-25T10:59:54+02:00 | GitHub | noreply@github.com | 2025-10-25T10:59:54+02:00 | | vulkan: delete dead code (#16732) |
| 248 | 8423d019318b446640bb620e4ce80066d8530f05 | 5cca2542ac3f3f86831d32bce744d08fc2b353b0 | Jeff Bolz | jbolz@nvidia.com | 2025-10-25T00:04:12-05:00 | GitHub | noreply@github.com | 2025-10-25T07:04:12+02:00 | | vulkan: Optimize SSM_SCAN (#16645) |
| 249 | 5cca2542ac3f3f86831d32bce744d08fc2b353b0 | 55945d2ef51b93821d4b6f4a9b994393344a90db | compilade | git@compilade.net | 2025-10-24T20:52:00-04:00 | GitHub | noreply@github.com | 2025-10-24T20:52:00-04:00 | | convert : avoid dequantizing mxfp4 for GPT-OSS (#16756) |
| 250 | 55945d2ef51b93821d4b6f4a9b994393344a90db | 0bcb40b48c6fc6f17ba9672625e526ab2574344b | leejet | leejet714@gmail.com | 2025-10-25T03:39:37+08:00 | GitHub | noreply@github.com | 2025-10-24T21:39:37+02:00 | | ggml: fix CUDA grid launch condition for large block_nums.y in binbcast (#16742) |
| 251 | 0bcb40b48c6fc6f17ba9672625e526ab2574344b | 69e9ff010309a1155d704cf9320bdb3aaf4160ca | Aman Gupta | amangupta052@gmail.com | 2025-10-24T20:46:19+08:00 | GitHub | noreply@github.com | 2025-10-24T20:46:19+08:00 | | CUDA: use CUB for arbitary size argsort (#16754) |
| 252 | 69e9ff010309a1155d704cf9320bdb3aaf4160ca | 5a91109a5d7dab5d7adc40bedb397ede99a705b1 | Florian Badie | florianbadie@odrling.xyz | 2025-10-24T14:10:29+02:00 | GitHub | noreply@github.com | 2025-10-24T14:10:29+02:00 | | webui: support q URL parameter (#16728) |
| 253 | 5a91109a5d7dab5d7adc40bedb397ede99a705b1 | f8f071faddf32ea09f4234edb6e809b380a9ee26 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-24T12:02:02+02:00 | GitHub | noreply@github.com | 2025-10-24T12:02:02+02:00 | | model-conversion : add trust_remote_code for orig model run [no ci] (#16751) |
| 254 | f8f071faddf32ea09f4234edb6e809b380a9ee26 | 0bf47a1dbba4d36f2aff4e8c34b06210ba34e688 | compilade | git@compilade.net | 2025-10-23T16:31:41-04:00 | GitHub | noreply@github.com | 2025-10-23T16:31:41-04:00 | | convert : handle pre-quantized models (#14810) |
| 255 | 0bf47a1dbba4d36f2aff4e8c34b06210ba34e688 | dd62dcfab97e420949519fd0eac9fca7bf97e635 | Johannes Gäßler | johannesg@5d6.de | 2025-10-23T21:30:17+02:00 | GitHub | noreply@github.com | 2025-10-23T21:30:17+02:00 | | server: add memory breakdown print (#16740) |
| 256 | dd62dcfab97e420949519fd0eac9fca7bf97e635 | d0660f237a5c31771a3d6d1030ebe3e0c409ba92 | Julien Denize | 40604584+juliendenize@users.noreply.github.com | 2025-10-23T15:54:46+02:00 | GitHub | noreply@github.com | 2025-10-23T15:54:46+02:00 | | convert : Make mistral-common dependency optional (#16738) |
| 257 | d0660f237a5c31771a3d6d1030ebe3e0c409ba92 | fe6a9882acf5c02f96624ed8f80144100d7006cb | Xuan-Son Nguyen | son@huggingface.co | 2025-10-23T15:00:49+02:00 | GitHub | noreply@github.com | 2025-10-23T15:00:49+02:00 | | mtmd-cli : allow using --jinja (#16718) |
| 258 | fe6a9882acf5c02f96624ed8f80144100d7006cb | 061f0eff02dc9a82f7bd850db3bd70b8a0b5e87a | Prajwal B Mehendarkar | prajwal.b.mehendarkar@ibm.com | 2025-10-23T17:07:31+05:30 | GitHub | noreply@github.com | 2025-10-23T19:37:31+08:00 | | Manually link -lbsd to resolve flock symbol on AIX (#16610) |
| 259 | 061f0eff02dc9a82f7bd850db3bd70b8a0b5e87a | 8cf6b42d467d05fa7d9776d2bcc69974ecce6900 | Aman Gupta | amangupta052@gmail.com | 2025-10-23T19:14:06+08:00 | GitHub | noreply@github.com | 2025-10-23T19:14:06+08:00 | | ggml-cuda: use passed ops instead of hardcoded ops (#16712) |
| 260 | 8cf6b42d467d05fa7d9776d2bcc69974ecce6900 | 9de9672adb0f4ca4e39483ac3ffed52b3f70a55d | matteo | matteo.serva@gmail.com | 2025-10-23T11:32:24+02:00 | GitHub | noreply@github.com | 2025-10-23T12:32:24+03:00 | | server : send partial stop string when <EOG> is reached (#15007) |
| 261 | 9de9672adb0f4ca4e39483ac3ffed52b3f70a55d | 63d2fc46e17a06be5b4b5823a5ada088317f1f0a | Matthew Michel | matthew.michel@intel.com | 2025-10-22T20:05:15-05:00 | GitHub | noreply@github.com | 2025-10-23T09:05:15+08:00 | | sycl: use async memory allocation to fix crashes during graph recording (#16644) |
| 262 | 63d2fc46e17a06be5b4b5823a5ada088317f1f0a | a2e0088d9242bd9e57f8b852b05a6e47843b5a45 | Max Krasnyansky | maxk@qti.qualcomm.com | 2025-10-22T13:47:09-07:00 | GitHub | noreply@github.com | 2025-10-22T13:47:09-07:00 | | Add experimental ggml-hexagon backend for the Hexagon NPU (#16547) |
| 263 | 9b9201f65a22c02cee8e300f58f480a588591227 | 19a5a3edfd306516cc419679d69d6435943b6816 | Pascal | admin@serveurperso.com | 2025-10-22T16:58:23+02:00 | GitHub | noreply@github.com | 2025-10-22T16:58:23+02:00 | | webui: introduce OpenAI-compatible model selector in JSON payload (#16562) |
| 264 | 19a5a3edfd306516cc419679d69d6435943b6816 | d8eaa26e4d9228df3aa46a930db60c8eaab67c1b | sirus20x6 | sirus20x6@users.noreply.github.com | 2025-10-22T05:14:14-05:00 | GitHub | noreply@github.com | 2025-10-22T12:14:14+02:00 | | ggml : Leverage the existing GGML_F32_VEC helpers to vectorize ggml_vec_set_f32 for faster fills (#16522) |
| 265 | d8eaa26e4d9228df3aa46a930db60c8eaab67c1b | 9285325ce0174631c3cd6121d56084adc4ef2d8f | Acly | aclysia@gmail.com | 2025-10-22T12:01:22+02:00 | GitHub | noreply@github.com | 2025-10-22T12:01:22+02:00 | | tests : fix test-thread-safety when compiling with multiple backends (#16699) |
| 266 | 9285325ce0174631c3cd6121d56084adc4ef2d8f | 03792ad93609fc67e41041c6347d9aa14e5e0d74 | Aman Gupta | amangupta052@gmail.com | 2025-10-22T12:33:08+08:00 | GitHub | noreply@github.com | 2025-10-22T12:33:08+08:00 | | CUDA: fix bug in topk-moe softmax (#16711) |
| 267 | 03792ad93609fc67e41041c6347d9aa14e5e0d74 | 51d1a8c997bd2629ef211a30208058ea87a30982 | Aman Gupta | amangupta052@gmail.com | 2025-10-21T22:40:38+08:00 | GitHub | noreply@github.com | 2025-10-21T22:40:38+08:00 | | CUDA: topk-moe: add optional parameter for gpt-oss (#16649) |
| 268 | 51d1a8c997bd2629ef211a30208058ea87a30982 | 4926419c4d74a1cf724e7163d937eb72f36e7b26 | Johannes Gäßler | johannesg@5d6.de | 2025-10-21T15:27:53+02:00 | GitHub | noreply@github.com | 2025-10-21T15:27:53+02:00 | | CUDA: better error for FA kernel with 0 occupancy (#16643) |
| 269 | 4926419c4d74a1cf724e7163d937eb72f36e7b26 | 6ea37f57391d27736c35cd3c20c1f990b7952b74 | Aman Gupta | amangupta052@gmail.com | 2025-10-21T16:43:14+08:00 | GitHub | noreply@github.com | 2025-10-21T16:43:14+08:00 | | ggml: add ggml_can_fuse_subgraph (#16662) |
| 270 | 6ea37f57391d27736c35cd3c20c1f990b7952b74 | fb349848f387f355450c3187556e71e6d32c145f | lhez | lih@qti.qualcomm.com | 2025-10-20T22:26:17-07:00 | GitHub | noreply@github.com | 2025-10-20T22:26:17-07:00 | | opencl: fix warnings and clean up profiling (#16688) |
| 271 | fb349848f387f355450c3187556e71e6d32c145f | 6de8ed75196c7cd98c1f34bbf3a7452451ba8ac2 | Jeff Bolz | jbolz@nvidia.com | 2025-10-20T22:16:08-05:00 | GitHub | noreply@github.com | 2025-10-20T22:16:08-05:00 | | vulkan: Handle FA with all -inf mask values (#16447) |
| 272 | 6de8ed75196c7cd98c1f34bbf3a7452451ba8ac2 | 84bf3c677857279037adf67cdcfd89eaa4ca9281 | YehuditE | y8703470@gmail.com | 2025-10-21T01:21:12+03:00 | GitHub | noreply@github.com | 2025-10-21T00:21:12+02:00 | | sycl : add PAD_REFLECT_D1 operator support (#16145) |
| 273 | 84bf3c677857279037adf67cdcfd89eaa4ca9281 | c9c1972e2c2cc6a771fcc145bfa138700179f961 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-20T21:38:20+02:00 | GitHub | noreply@github.com | 2025-10-20T21:38:20+02:00 | | model : add BailingMoeV2 support (#16063) |
| 274 | c9c1972e2c2cc6a771fcc145bfa138700179f961 | b617cfd2896edd592a36ebbc041817eb030a1005 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-20T19:49:02+02:00 | GitHub | noreply@github.com | 2025-10-20T19:49:02+02:00 | | Handle legacy 'context' attachments (#16687) |
| 275 | b617cfd2896edd592a36ebbc041817eb030a1005 | 79068501fac9a74cca7129a8e5a8281b410a853e | Diego Devesa | slarengh@gmail.com | 2025-10-20T05:53:50-07:00 | GitHub | noreply@github.com | 2025-10-20T14:53:50+02:00 | | ggml-alloc : fix leak when reusing a tensor with a larger size (#16679) |
| 276 | 79068501fac9a74cca7129a8e5a8281b410a853e | 0e4a0cf2fae667d3efcf52f2f52398779d986b1d | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-20T14:21:12+02:00 | GitHub | noreply@github.com | 2025-10-20T14:21:12+02:00 | | Prevent premature submission on IME input (#16673) |
| 277 | 0e4a0cf2fae667d3efcf52f2f52398779d986b1d | 13f2cfad4170c096c51a02c24a6a158cb47f1480 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-20T13:29:14+02:00 | GitHub | noreply@github.com | 2025-10-20T13:29:14+02:00 | | Import/Export UX improvements (#16619) |
| 278 | 13f2cfad4170c096c51a02c24a6a158cb47f1480 | 06332e28672356b964d6dfc2ba4657e20581cd43 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-20T12:41:13+02:00 | GitHub | noreply@github.com | 2025-10-20T12:41:13+02:00 | | Enable per-conversation loading states to allow having parallel conversations (#16327) |
| 279 | 06332e28672356b964d6dfc2ba4657e20581cd43 | 72d53e6c4decee8b339e49aed8cc0e234b9639dc | takuya kodama | otegami@clear-code.com | 2025-10-20T16:27:09+08:00 | GitHub | noreply@github.com | 2025-10-20T11:27:09+03:00 | | llama-batch: fix build fails with `-Werror=missing-braces` (#16614) |
| 280 | 72d53e6c4decee8b339e49aed8cc0e234b9639dc | 2330de7b847ca84eac766df372c604c26db72747 | Ron Evans | ron@hybridgroup.com | 2025-10-20T10:20:04+02:00 | GitHub | noreply@github.com | 2025-10-20T11:20:04+03:00 | | readme: update bindings (#16651) |
| 281 | 2330de7b847ca84eac766df372c604c26db72747 | 7062dd8460685d6700ed7621e50a22c6f3400ca3 | safranowith | bsh155762@gmail.com | 2025-10-20T11:08:32+03:00 | GitHub | noreply@github.com | 2025-10-20T11:08:32+03:00 | | SYCL: Add support for FLOOR,CEIL,ROUND and TRUNC unary operators (#16613) |
| 282 | 7062dd8460685d6700ed7621e50a22c6f3400ca3 | 0398752dd450dfabdd1b9e289f6364c2600f6ab5 | takuya kodama | otegami@clear-code.com | 2025-10-20T15:44:21+08:00 | GitHub | noreply@github.com | 2025-10-20T10:44:21+03:00 | | llama-context: only warn on pooling_type when user specified (#16674) |
| 283 | c20c0a5c0c44179f5267a262546624854b1203d1 | 62425b3a7e8b169a3bbe6ce8c7324513a7fcee8c | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-20T14:41:08+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-20T14:41:08+08:00 | | kv: turn on common-prefix mtmd: fix output size bug |
| 284 | 0398752dd450dfabdd1b9e289f6364c2600f6ab5 | 4f73d0a95120687e2c527739f771330a5271259a | Giuseppe Scrivano | gscrivan@redhat.com | 2025-10-19T23:54:31+02:00 | GitHub | noreply@github.com | 2025-10-19T23:54:31+02:00 | | model : add Granite Hybrid types (#16635) |
| 285 | 4f73d0a95120687e2c527739f771330a5271259a | cec5edbcaec69bbf6d5851cabce4ac148be41701 | Aaron Teo | aaron.teo1@ibm.com | 2025-10-20T05:06:39+08:00 | GitHub | noreply@github.com | 2025-10-19T23:06:39+02:00 | | ci : fix binaries release failure for s390x (binaries may not work yet) (#16664) |
| 286 | cec5edbcaec69bbf6d5851cabce4ac148be41701 | fcb235b46618921cbd826acd49b553b5302233aa | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-19T14:03:25+02:00 | GitHub | noreply@github.com | 2025-10-19T14:03:25+02:00 | | ci : avoid manual updates of docs/ops.md (#16663) |
| 287 | fcb235b46618921cbd826acd49b553b5302233aa | 55754bebd5d570960cde9c0ba991a0b2991f6b1e | Aaron Teo | aaron.teo1@ibm.com | 2025-10-19T18:37:47+08:00 | GitHub | noreply@github.com | 2025-10-19T18:37:47+08:00 | | ci: include s390x release binaries (#16648) |
| 288 | 55754bebd5d570960cde9c0ba991a0b2991f6b1e | ee09828cb057460b369576410601a3a09279e23c | Aman Gupta | amangupta052@gmail.com | 2025-10-19T15:37:12+08:00 | GitHub | noreply@github.com | 2025-10-19T10:37:12+03:00 | | CODEOWNERS: update for ggml-cuda/mmf (#16660) |
| 289 | ee09828cb057460b369576410601a3a09279e23c | e56abd2098dd2e2b0804691b93c13b48ae421627 | Johannes Gäßler | johannesg@5d6.de | 2025-10-18T14:47:32+02:00 | GitHub | noreply@github.com | 2025-10-18T14:47:32+02:00 | | HIP: fix GPU_TARGETS (#16642) |
| 290 | e56abd2098dd2e2b0804691b93c13b48ae421627 | 38355c6c8e43204e11a22daa7483082c0ff01e71 | Jeff Bolz | jbolz@nvidia.com | 2025-10-18T05:22:57-05:00 | GitHub | noreply@github.com | 2025-10-18T12:22:57+02:00 | | vulkan: Implement topk_moe fused shader, ported from CUDA (#16641) |
| 291 | 38355c6c8e43204e11a22daa7483082c0ff01e71 | 81387858f1fbcc1acedbd308486e1016618ca8f8 | Aman Gupta | amangupta052@gmail.com | 2025-10-18T17:52:53+08:00 | GitHub | noreply@github.com | 2025-10-18T11:52:53+02:00 | | CUDA: use registers instead of smem in topk-moe (#16647) |
| 292 | 81387858f1fbcc1acedbd308486e1016618ca8f8 | 66b0dbcb2d462e7b70ba5a69ee8c3899ac2efb1c | Shawn Gu | shawngu@qti.qualcomm.com | 2025-10-17T17:55:32-07:00 | GitHub | noreply@github.com | 2025-10-17T17:55:32-07:00 | | opencl: transposed gemm/gemv moe kernel with mxfp4,f32 (#16602) |
| 293 | 66b0dbcb2d462e7b70ba5a69ee8c3899ac2efb1c | 41386cf365d894134ee0813d15e2f5d76f6a4d8e | Johannes Gäßler | johannesg@5d6.de | 2025-10-17T17:41:09+02:00 | GitHub | noreply@github.com | 2025-10-17T17:41:09+02:00 | | llama-model: fix insonsistent ctxs <-> bufs order (#16581) |
| 294 | 41386cf365d894134ee0813d15e2f5d76f6a4d8e | 3d4e86bbeb15f487d6da6174ba6191b7c212cc25 | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-17T18:02:52+03:00 | GitHub | noreply@github.com | 2025-10-17T18:02:52+03:00 | | rpc : report actual free memory (#16616) |
| 295 | 3d4e86bbeb15f487d6da6174ba6191b7c212cc25 | 342c728d031d50673feded797520a44127d73379 | Giuseppe Scrivano | gscrivan@redhat.com | 2025-10-17T14:23:47+02:00 | GitHub | noreply@github.com | 2025-10-17T14:23:47+02:00 | | vulkan: Add State Space Model (SSM) Operations Support (#16463) |
| 296 | 342c728d031d50673feded797520a44127d73379 | ababae7e1ec3e9cfdab0322ee55ea3389e82a4d5 | muggle-stack | promuggle@qq.com | 2025-10-17T18:01:23+08:00 | GitHub | noreply@github.com | 2025-10-17T13:01:23+03:00 | | ggml : fix SpaceMit IME array out-of-bounds in task assignment (#16629) |
| 297 | ababae7e1ec3e9cfdab0322ee55ea3389e82a4d5 | b19491599d4d42c606601d75e95c1d1de3291f8e | Pascal | admin@serveurperso.com | 2025-10-17T10:35:03+02:00 | GitHub | noreply@github.com | 2025-10-17T10:35:03+02:00 | | webui: reorganize settings layout (#16607) |
| 298 | b19491599d4d42c606601d75e95c1d1de3291f8e | 9ad4f1931ee0f3b41d9355245ef744786aaae0aa | Jeff Bolz | jbolz@nvidia.com | 2025-10-17T02:31:04-05:00 | GitHub | noreply@github.com | 2025-10-17T09:31:04+02:00 | | vulkan: fix debug build (add_rms_len/data not found) (#16624) |
| 299 | 9ad4f1931ee0f3b41d9355245ef744786aaae0aa | 79967ec596c0dacfd2251b085a57e79df292b1cc | Ilia Ilmer | iliailmer@users.noreply.github.com | 2025-10-17T02:33:58-04:00 | GitHub | noreply@github.com | 2025-10-17T09:33:58+03:00 | | metal : add `CONV_TRANSPOSE_2D` (#16542) |
| 300 | 79967ec596c0dacfd2251b085a57e79df292b1cc | ceff6bb253dd306f5404d7ccb3f11fadafe71b52 | Olivier Chafik | olivier.chafik@gmail.com | 2025-10-17T06:59:31+01:00 | GitHub | noreply@github.com | 2025-10-17T08:59:31+03:00 | | grammar : use int64_t to avoid int overflows in int schema to grammar conversion logic (#16626) |
| 301 | ceff6bb253dd306f5404d7ccb3f11fadafe71b52 | 1bb4f43380944e94c9a86e305789ba103f5e62bd | GittyBurstein | g0534163997@gmail.com | 2025-10-17T05:36:40+03:00 | GitHub | noreply@github.com | 2025-10-17T10:36:40+08:00 | | SYCL SET operator optimized for F32 tensors (#16350) |
| 302 | 1bb4f43380944e94c9a86e305789ba103f5e62bd | 683fa6ba4ed3a23b939d3e11e6dc860bd47a0ccf | Xuan-Son Nguyen | son@huggingface.co | 2025-10-16T19:00:31+02:00 | GitHub | noreply@github.com | 2025-10-16T19:00:31+02:00 | | mtmd : support home-cooked Mistral Small Omni (#14928) |
| 303 | 683fa6ba4ed3a23b939d3e11e6dc860bd47a0ccf | b22572e97dc51757d3ebe917a5a283385010ec68 | Pascal | admin@serveurperso.com | 2025-10-16T16:28:41+02:00 | GitHub | noreply@github.com | 2025-10-16T16:28:41+02:00 | | fix: added a normalization step for MathJax-style \[\] and \(\) delimiters (#16599) |
| 304 | b22572e97dc51757d3ebe917a5a283385010ec68 | 7a50cf388a530127cbbdd2a507ef81a451c9d819 | GittyBurstein | g0534163997@gmail.com | 2025-10-16T16:26:21+03:00 | GitHub | noreply@github.com | 2025-10-16T15:26:21+02:00 | | sycl : add ARANGE operator (#16362) |
| 305 | 62425b3a7e8b169a3bbe6ce8c7324513a7fcee8c | 76a63a2998286c1649de2496bccb7d8a46549895 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-16T17:34:39+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-16T17:34:39+08:00 | | kv: let seq_rm free whole seq since current hw-op cannot support common-prefix |
| 306 | 7a50cf388a530127cbbdd2a507ef81a451c9d819 | 6f5d924637a15abedb111cbbffd7da5f31c81855 | Chenguang Li | 757486878@qq.com | 2025-10-16T16:41:11+08:00 | GitHub | noreply@github.com | 2025-10-16T16:41:11+08:00 | | CANN: format code using .clang-format (#15863) |
| 307 | 76a63a2998286c1649de2496bccb7d8a46549895 | b67217078aa748b06e1b2269d88dd18c0aca2c26 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-16T16:31:15+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-16T16:31:15+08:00 | | kv: fix common-prefix bug by implementing llama_memory_seq_rm |
| 308 | 6f5d924637a15abedb111cbbffd7da5f31c81855 | adc9b60f190c1016a09f439862fa1cbb302262ac | takasurazeem | takasurazeem@gmail.com | 2025-10-16T01:11:33-04:00 | GitHub | noreply@github.com | 2025-10-16T08:11:33+03:00 | | common : Update the docs on -t --threads (#16236) |
| 309 | adc9b60f190c1016a09f439862fa1cbb302262ac | ee50ee1eadff58777ae746827b04de7ba0befc55 | takuya kodama | a.s.takuya1026@gmail.com | 2025-10-16T13:10:32+08:00 | GitHub | noreply@github.com | 2025-10-16T08:10:32+03:00 | | ggml-cpu: replace putenv with setenv for const-correctness (#16573) |
| 310 | ee50ee1eadff58777ae746827b04de7ba0befc55 | 7adc79c03234de9a20661fd6dbf2d02c32ca7acb | yael-works | 106673277+yael-works@users.noreply.github.com | 2025-10-16T07:21:28+03:00 | GitHub | noreply@github.com | 2025-10-16T12:21:28+08:00 | | SYCL: Add GGML_OP_MEAN operator support (#16009) |
| 311 | 7adc79c03234de9a20661fd6dbf2d02c32ca7acb | 466c1911ab736f0b7366127edee99f8ee5687417 | Aleksei Nikiforov | 103434461+AlekseiNikiforovIBM@users.noreply.github.com | 2025-10-15T22:43:08+02:00 | GitHub | noreply@github.com | 2025-10-15T22:43:08+02:00 | | gguf-py : add support for endian conversion of BF16 data (#16594) |
| 312 | 466c1911ab736f0b7366127edee99f8ee5687417 | 0cb7a0683b0529172472d74d21f05470a607f297 | safranowith | bsh155762@gmail.com | 2025-10-15T22:24:51+03:00 | GitHub | noreply@github.com | 2025-10-15T21:24:51+02:00 | | cpu : add FLOOR, CEIL, ROUND and TRUNC unary operators (#16083) |
| 313 | 0cb7a0683b0529172472d74d21f05470a607f297 | d93f8439b08c4f35e13a41a7366901fdbe770fc8 | lhez | lih@qti.qualcomm.com | 2025-10-15T10:51:04-07:00 | GitHub | noreply@github.com | 2025-10-15T10:51:04-07:00 | | opencl: add q8_0 mm support (#16469) |
| 314 | d93f8439b08c4f35e13a41a7366901fdbe770fc8 | f9fb33f2630b4b4ba9081ce9c0c921f8cd8ba4eb | lhez | lih@qti.qualcomm.com | 2025-10-15T10:48:28-07:00 | GitHub | noreply@github.com | 2025-10-15T10:48:28-07:00 | | opencl: fix FA for f32 (#16584) |
| 315 | f9fb33f2630b4b4ba9081ce9c0c921f8cd8ba4eb | f4ce81c45e7bd910e36bf44c253fc5255c49b1e4 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-15T16:22:20+02:00 | GitHub | noreply@github.com | 2025-10-15T16:22:20+02:00 | | Add server-driven parameter defaults and syncing (#16515) |
| 316 | f4ce81c45e7bd910e36bf44c253fc5255c49b1e4 | 17304cbcc1dd24de7741cbe57925d58e90a98ac1 | Sam/Samuel | 57896620+cern1710@users.noreply.github.com | 2025-10-15T23:05:56+09:00 | GitHub | noreply@github.com | 2025-10-15T17:05:56+03:00 | | metal: optimise `GGML_OP_SUM` (#16559) |
| 317 | 17304cbcc1dd24de7741cbe57925d58e90a98ac1 | 3e3cb19f6449f9168a128eaeae01f8f41b049acc | Georgi Gerganov | ggerganov@gmail.com | 2025-10-15T16:53:12+03:00 | GitHub | noreply@github.com | 2025-10-15T16:53:12+03:00 | | server : fix img token logs (#16595) |
| 318 | 3e3cb19f6449f9168a128eaeae01f8f41b049acc | 5acd455460f457942d8dd02e3dd9b1eebfce99fe | Xuan-Son Nguyen | son@huggingface.co | 2025-10-15T14:48:08+02:00 | GitHub | noreply@github.com | 2025-10-15T14:48:08+02:00 | | llama-quant: add support for mmproj (#16592) |
| 319 | 5acd455460f457942d8dd02e3dd9b1eebfce99fe | 554fd578a5ed78a10f371f5850c6a69aa83df15a | Julius Tischbein | ju.tischbein@gmail.com | 2025-10-15T13:54:15+02:00 | GitHub | noreply@github.com | 2025-10-15T14:54:15+03:00 | | CUDA: Changing the CUDA scheduling strategy to spin (#16585) |
| 320 | 554fd578a5ed78a10f371f5850c6a69aa83df15a | fa882fd2b1bcb663de23af06fdc391489d05b007 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-15T12:51:27+03:00 | GitHub | noreply@github.com | 2025-10-15T11:51:27+02:00 | | server : fix mtmd checkpoints (#16591) |
| 321 | fa882fd2b1bcb663de23af06fdc391489d05b007 | ffa059034c1e41f3e58363ef3c46d2fcf854bc4d | Georgi Gerganov | ggerganov@gmail.com | 2025-10-14T20:33:05+03:00 | GitHub | noreply@github.com | 2025-10-14T20:33:05+03:00 | | metal : avoid using Metal's gpuAddress property (#16576) |
| 322 | ffa059034c1e41f3e58363ef3c46d2fcf854bc4d | 120bf7046d85a893f064c12abe58bfeebd735f84 | SavicStefan | 50296686+SavicStefan@users.noreply.github.com | 2025-10-14T19:18:05+02:00 | GitHub | noreply@github.com | 2025-10-14T19:18:05+02:00 | | vulkan: Add ACC_TYPE_VEC2 implementation (#16203) |
| 323 | 120bf7046d85a893f064c12abe58bfeebd735f84 | 4258e0cfe72c72ad2787be1e9f93253e89219455 | Aman Gupta | amangupta052@gmail.com | 2025-10-14T22:48:08+08:00 | GitHub | noreply@github.com | 2025-10-14T07:48:08-07:00 | | CUDA + openCL: fix bug in accessing rms_norm->src while doing fusion (#16577) |
| 324 | 4258e0cfe72c72ad2787be1e9f93253e89219455 | 7ea15bb64c81e3813eb0babf9a57e1bc5697f569 | Jeff Bolz | jbolz@nvidia.com | 2025-10-14T08:53:37-05:00 | GitHub | noreply@github.com | 2025-10-14T15:53:37+02:00 | | vulkan: Support FA with K/V in F32 (#16543) |
| 325 | 7ea15bb64c81e3813eb0babf9a57e1bc5697f569 | 9c7185dd28416cf67f5e3b268381f311b5e3da56 | Jeff Bolz | jbolz@nvidia.com | 2025-10-14T07:51:36-05:00 | GitHub | noreply@github.com | 2025-10-14T14:51:36+02:00 | | vulkan: Improve build time for MSVC (#16545) |
| 326 | 9c7185dd28416cf67f5e3b268381f311b5e3da56 | 1ee9d0b415cdf5240418c110a18b419f4002b154 | Johannes Gäßler | johannesg@5d6.de | 2025-10-14T14:22:47+02:00 | GitHub | noreply@github.com | 2025-10-14T14:22:47+02:00 | | CUDA: enable FA for FP32 KV cache (#16546) |
| 327 | 1ee9d0b415cdf5240418c110a18b419f4002b154 | 48e2fa9fb7c2de1e53808fdb65ec33f916020fc4 | Aman Gupta | amangupta052@gmail.com | 2025-10-14T19:16:21+08:00 | GitHub | noreply@github.com | 2025-10-14T13:16:21+02:00 | | CUDA: use fastdiv + ggml_cuda_mad for mmvf (#16557) |
| 328 | 48e2fa9fb7c2de1e53808fdb65ec33f916020fc4 | 5b6913c47b6bc71a6f927805a45387d5657d8b89 | Aman Gupta | amangupta052@gmail.com | 2025-10-14T19:15:15+08:00 | GitHub | noreply@github.com | 2025-10-14T13:15:15+02:00 | | CUDA: add fp kernel for larger batch size MoE (#16512) |
| 329 | 5b6913c47b6bc71a6f927805a45387d5657d8b89 | bc07349a7f87ba6eb31ed4b0ea9d9a7352185213 | Anav Prasad | anavp@nvidia.com | 2025-10-14T09:53:49Z | GitHub | noreply@github.com | 2025-10-14T11:53:49+02:00 | | cuda : remove legacy copy-op pointer indirection code (#16485) |
| 330 | bc07349a7f87ba6eb31ed4b0ea9d9a7352185213 | e60f241eacec42d3bd7c9edd37d236ebf35132a8 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-14T08:48:50+03:00 | GitHub | noreply@github.com | 2025-10-14T08:48:50+03:00 | | server : dynamic token limit for prompt cache (#16560) |
| 331 | e60f241eacec42d3bd7c9edd37d236ebf35132a8 | e38b7c6e9e4453e3b3e96d76e38bc2ccb6bce458 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-13T23:07:57+03:00 | GitHub | noreply@github.com | 2025-10-13T23:07:57+03:00 | | metal : FA support F32 K and V and head size = 32 (#16531) |
| 332 | e38b7c6e9e4453e3b3e96d76e38bc2ccb6bce458 | 5016b7286240d29f8f640039989b84ea3a854344 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-13T22:42:37+03:00 | GitHub | noreply@github.com | 2025-10-13T22:42:37+03:00 | | graph : support cacheless embeddings with FA and iSWA (#16528) |
| 333 | 5016b7286240d29f8f640039989b84ea3a854344 | 7049736b2dd9011bf819e298b844ebbc4b5afdc9 | lhez | lih@qti.qualcomm.com | 2025-10-13T11:50:37-07:00 | GitHub | noreply@github.com | 2025-10-13T11:50:37-07:00 | | opencl: fix build targeting CL 2 (#16554) |
| 334 | 7049736b2dd9011bf819e298b844ebbc4b5afdc9 | 01d2bdc2bc61b676706830305b286b08b9885a41 | Johannes Gäßler | johannesg@5d6.de | 2025-10-13T16:29:45+02:00 | GitHub | noreply@github.com | 2025-10-13T17:29:45+03:00 | | CUDA: fix numerical issues in tile FA kernel (#16540) |
| 335 | 01d2bdc2bc61b676706830305b286b08b9885a41 | 56fc38b9655fbe1869d8bd6cfb269418196cea69 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-10-13T20:48:47+08:00 | GitHub | noreply@github.com | 2025-10-13T15:48:47+03:00 | | ggml : fix build broken with -march=armv9-a on MacOS (#16520) |
| 336 | 56fc38b9655fbe1869d8bd6cfb269418196cea69 | 1fb9504eb744969a990bfe4cfcf1d3d7a479541c | Chenguang Li | 757486878@qq.com | 2025-10-13T17:01:24+08:00 | GitHub | noreply@github.com | 2025-10-13T17:01:24+08:00 | | CANN: fix CPU memory leak in CANN backend (#16549) |
| 337 | 1fb9504eb744969a990bfe4cfcf1d3d7a479541c | 3f750f8d760ab5a61491e6a9409072dfeee4b4d7 | Pascal | admin@serveurperso.com | 2025-10-13T10:55:32+02:00 | GitHub | noreply@github.com | 2025-10-13T10:55:32+02:00 | | fix: add remark plugin to render raw HTML as literal text (#16505) |
| 338 | b67217078aa748b06e1b2269d88dd18c0aca2c26 | 98c8269abefdfa8fde47dfb2769ceebc8b704e86 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-13T11:48:14+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-10-13T16:52:48+08:00 | | add different path for unified and seperate compilation |
| 339 | 3f750f8d760ab5a61491e6a9409072dfeee4b4d7 | c515fc577166042234241c6bd0da9b08dcbe2bb9 | Sam/Samuel | 57896620+cern1710@users.noreply.github.com | 2025-10-13T16:25:02+08:00 | GitHub | noreply@github.com | 2025-10-13T11:25:02+03:00 | | metal: add support for opt_step_sgd (#16539) |
| 340 | c515fc577166042234241c6bd0da9b08dcbe2bb9 | f9bc66c3ebcfddb5f09e4b21253623caeb8e414a | Georgi Gerganov | ggerganov@gmail.com | 2025-10-13T11:22:27+03:00 | GitHub | noreply@github.com | 2025-10-13T11:22:27+03:00 | | ggml : fix scalar path for computing norm (#16558) |
| 341 | f9bc66c3ebcfddb5f09e4b21253623caeb8e414a | a31cf36ad946a13b3a646bf0dadf2a481e89f944 | hipudding | huafengchun@gmail.com | 2025-10-13T08:52:22+08:00 | GitHub | noreply@github.com | 2025-10-13T08:52:22+08:00 | | CANN: Update several operators to support FP16 data format (#16251) |
| 342 | a31cf36ad946a13b3a646bf0dadf2a481e89f944 | 81d54bbfd599811b354c39f04550888168be7780 | Sam/Samuel | 57896620+cern1710@users.noreply.github.com | 2025-10-13T02:43:14+08:00 | GitHub | noreply@github.com | 2025-10-12T21:43:14+03:00 | | metal : add opt_step_adamw and op_sum (#16529) |
| 343 | 81d54bbfd599811b354c39f04550888168be7780 | c7be9febcbafa9af7d1b9443f86475c59c9c5f87 | Pascal | admin@serveurperso.com | 2025-10-12T18:06:41+02:00 | GitHub | noreply@github.com | 2025-10-12T18:06:41+02:00 | | webui: remove client-side context pre-check and rely on backend for limits (#16506) |
| 344 | c7be9febcbafa9af7d1b9443f86475c59c9c5f87 | 8415f61e23d04427cd0d912fbb9d33b85f849456 | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-10-12T21:53:35+08:00 | GitHub | noreply@github.com | 2025-10-12T21:53:35+08:00 | | [SYCL] fix UT fault cases: count-equal, argsort, pad OPs (#16521) |
| 345 | 8415f61e23d04427cd0d912fbb9d33b85f849456 | 2c301e91abb92d03c1a682b4b540ba835562a74b | Mathieu Baudier | mbaudier@argeo.org | 2025-10-12T15:48:03+02:00 | GitHub | noreply@github.com | 2025-10-12T15:48:03+02:00 | | ci : add Vulkan on Ubuntu with default packages build (#16532) |
| 346 | 2c301e91abb92d03c1a682b4b540ba835562a74b | 4b2dae383df708e2afc49c4859a81cd074f5ac10 | Aldehir Rojas | hello@alde.dev | 2025-10-12T08:18:47-05:00 | GitHub | noreply@github.com | 2025-10-12T16:18:47+03:00 | | common : handle unicode during partial json parsing (#16526) |
| 347 | 4b2dae383df708e2afc49c4859a81cd074f5ac10 | 41aac5c69b5fb281bc1f486afb053f78101bb39e | Georgi Gerganov | ggerganov@gmail.com | 2025-10-12T09:29:13+03:00 | GitHub | noreply@github.com | 2025-10-12T09:29:13+03:00 | | common : update presets (#16504) |
| 348 | 41aac5c69b5fb281bc1f486afb053f78101bb39e | a2fba89a426ff8005d303c73f0436e7e67368b70 | sirus20x6 | sirus20x6@users.noreply.github.com | 2025-10-12T00:25:37-05:00 | GitHub | noreply@github.com | 2025-10-12T08:25:37+03:00 | | ggml : Fix FP16 ELU positive branch (#16519) |
| 349 | a2fba89a426ff8005d303c73f0436e7e67368b70 | 20cc625edc2264aae2779e71bef1593e6a4e8c43 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-12T07:19:06+02:00 | GitHub | noreply@github.com | 2025-10-12T07:19:06+02:00 | | hparams : add check for layer index in is_recurrent (#16511) |
| 350 | 20cc625edc2264aae2779e71bef1593e6a4e8c43 | 11f0af5504252e453d57406a935480c909e3ff37 | sirus20x6 | sirus20x6@users.noreply.github.com | 2025-10-12T00:15:00-05:00 | GitHub | noreply@github.com | 2025-10-12T08:15:00+03:00 | | ggml: Correct SVE implementation in ggml_vec_dot_f16_unroll (#16518) |
| 351 | 11f0af5504252e453d57406a935480c909e3ff37 | a3cb04744fb5c591985f53b749fef5407d07a145 | Johannes Gäßler | johannesg@5d6.de | 2025-10-11T20:54:32+02:00 | GitHub | noreply@github.com | 2025-10-11T20:54:32+02:00 | | CUDA: faster tile FA, add oob checks, more HSs (#16492) |
| 352 | a3cb04744fb5c591985f53b749fef5407d07a145 | 4a8fbe0a5eb75f339782ebcf29c17848122184d3 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-11T16:54:10+03:00 | GitHub | noreply@github.com | 2025-10-11T16:54:10+03:00 | | metal : fix mul-mm condition + fix mul-mv permuted kernels (#16494) |
| 353 | 4a8fbe0a5eb75f339782ebcf29c17848122184d3 | 31d0ff1869aa2ea31f4e96d5877d0343e9a2171b | Pascal | admin@serveurperso.com | 2025-10-11T15:50:49+02:00 | GitHub | noreply@github.com | 2025-10-11T15:50:49+02:00 | | feat: render user content as markdown option (#16358) |
| 354 | 31d0ff1869aa2ea31f4e96d5877d0343e9a2171b | 97870e64975b26c5e06a3540a8dc0ff601351e86 | Yann Follet | 131855179+YannFollet@users.noreply.github.com | 2025-10-11T21:39:04+08:00 | GitHub | noreply@github.com | 2025-10-11T16:39:04+03:00 | | server / ranking : add sorting and management of top_n (#16403) |
| 355 | 97870e64975b26c5e06a3540a8dc0ff601351e86 | 477a66b03501cf3bd067f8968b77ca4d053ff1bd | Diego Devesa | slarengh@gmail.com | 2025-10-11T04:02:26-07:00 | GitHub | noreply@github.com | 2025-10-11T13:02:26+02:00 | | cuda : avoid initializing unused devices (#16510) |
| 356 | 477a66b03501cf3bd067f8968b77ca4d053ff1bd | e60f01d941bc5b7fae62dd57fee4cec76ec0ea6e | amirai21 | 89905406+amirai21@users.noreply.github.com | 2025-10-11T11:33:41+03:00 | GitHub | noreply@github.com | 2025-10-11T10:33:41+02:00 | | convert : correctly handle LLaMA tokenizer for Jamba (#16470) |
| 357 | e60f01d941bc5b7fae62dd57fee4cec76ec0ea6e | 81086cd6a3ca1252f0dc0f938171648399179c53 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-10T22:15:05+03:00 | GitHub | noreply@github.com | 2025-10-10T22:15:05+03:00 | | server : fix division by zero when reporting stats (#16501) |
| 358 | 81086cd6a3ca1252f0dc0f938171648399179c53 | 68ee98ae181a5c83a5cc6261daeee69a1f588c15 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-10T17:17:31+03:00 | GitHub | noreply@github.com | 2025-10-10T17:17:31+03:00 | | vocab : mark EOT token for Granite models (#16499) |
| 359 | 68ee98ae181a5c83a5cc6261daeee69a1f588c15 | cdb6da468cc33323955a523738d2e1675aeb5e9a | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-10T17:11:07+03:00 | GitHub | noreply@github.com | 2025-10-10T16:11:07+02:00 | | server : return HTTP 400 if prompt exceeds context length (#16486) |
| 360 | cdb6da468cc33323955a523738d2e1675aeb5e9a | 6d69ab3f262eca3f0c6ad0ce075d2c80b9924d0e | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-10T13:22:27+03:00 | GitHub | noreply@github.com | 2025-10-10T13:22:27+03:00 | | server : log requests to /v1/completions (#16495) |
| 361 | 6d69ab3f262eca3f0c6ad0ce075d2c80b9924d0e | 1faa13a1187051af66b0fd9f0d6effe4c77f0b3e | Prajwal B Mehendarkar | prajwal.b.mehendarkar@ibm.com | 2025-10-10T13:45:46+05:30 | GitHub | noreply@github.com | 2025-10-10T11:15:46+03:00 | | cmake : Dont define XOPENSOURCE on AIX (#16481) |
| 362 | 1faa13a1187051af66b0fd9f0d6effe4c77f0b3e | 1deee0f8d494981c32597dca8b5f8696d399b0f2 | Pascal | admin@serveurperso.com | 2025-10-09T22:54:57+02:00 | GitHub | noreply@github.com | 2025-10-09T22:54:57+02:00 | | webui: updated the chat service to only include max_tokens in the req… (#16489) |
| 363 | 1deee0f8d494981c32597dca8b5f8696d399b0f2 | d00cbea63c671cd85a57adaa50abf60b3b87d86f | duduta | simona.gherman@gmail.com | 2025-10-09T22:11:15+03:00 | GitHub | noreply@github.com | 2025-10-09T21:11:15+02:00 | | cpu : optimize the ggml NORM operation (#15953) |
| 364 | d00cbea63c671cd85a57adaa50abf60b3b87d86f | 8328fd4bae76fc027f8ca0e9a05acd3788dabe3f | Georgi Gerganov | ggerganov@gmail.com | 2025-10-09T18:54:51+03:00 | GitHub | noreply@github.com | 2025-10-09T18:54:51+03:00 | | server : host-memory prompt caching (#16391) |
| 365 | 8328fd4bae76fc027f8ca0e9a05acd3788dabe3f | 56b4795842d852152222ca7a2d4304008facf1b9 | Pascal | admin@serveurperso.com | 2025-10-09T17:36:29+02:00 | GitHub | noreply@github.com | 2025-10-09T17:36:29+02:00 | | No markdown in cot (#16483) |
| 366 | 56b4795842d852152222ca7a2d4304008facf1b9 | 2c0d875ae6c6043fbbafd2831e160955e9ca6af1 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-09T14:35:22+02:00 | GitHub | noreply@github.com | 2025-10-09T14:35:22+02:00 | | model-conversion : add support for SentenceTransformers (#16387) |
| 367 | 2c0d875ae6c6043fbbafd2831e160955e9ca6af1 | aa4711d369b0d4adb6802b98ad6ea362f838710e | sudhiarm | sudhi.sathyavathy@arm.com | 2025-10-09T09:13:18+01:00 | GitHub | noreply@github.com | 2025-10-09T10:13:18+02:00 | | ci: add ARM64 Kleidiai build and test support (#16462) |
| 368 | aa4711d369b0d4adb6802b98ad6ea362f838710e | d80d6d2400b8faa80654c723cb5bdf9fc8f4db06 | Chenguang Li | 757486878@qq.com | 2025-10-09T15:50:25+08:00 | GitHub | noreply@github.com | 2025-10-09T15:50:25+08:00 | | CANN: Improve ACL graph matching (#16166) |
| 369 | d80d6d2400b8faa80654c723cb5bdf9fc8f4db06 | b2602137557b2b28a39e03612717d85ead9a6f5a | Charles Xu | charles.xu@arm.com | 2025-10-09T09:29:17+02:00 | GitHub | noreply@github.com | 2025-10-09T10:29:17+03:00 | | kleidiai: kernel interface refactoring (#16460) |
| 370 | b2602137557b2b28a39e03612717d85ead9a6f5a | e08db4259521de493b7aeb49dadf29ebd1ee966a | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-10-09T15:25:11+08:00 | GitHub | noreply@github.com | 2025-10-09T10:25:11+03:00 | | [SYCL] refactor soft_max, add soft_max_back (#16472) |
| 371 | e08db4259521de493b7aeb49dadf29ebd1ee966a | 12bbc3fa50b6df03318a4451c9a2210200a0b28d | Saba Fallah | 10401143+sfallah@users.noreply.github.com | 2025-10-09T08:39:18+02:00 | GitHub | noreply@github.com | 2025-10-09T09:39:18+03:00 | | model: EmbeddingGemma Adding Support for SentenceTransformers Dense Modules (#16367) |
| 372 | 12bbc3fa50b6df03318a4451c9a2210200a0b28d | 9d0882840e6c3fb62965d03af0e22880ea90e012 | Pascal | admin@serveurperso.com | 2025-10-08T22:18:41+02:00 | GitHub | noreply@github.com | 2025-10-08T23:18:41+03:00 | | refactor: centralize CoT parsing in backend for streaming mode (#16394) |
| 373 | 9d0882840e6c3fb62965d03af0e22880ea90e012 | d2ee056e1df0fdf646270fb5621e9f92084b59a7 | ai-fonsi | length-amiss-7k@icloud.com | 2025-10-08T20:21:46+02:00 | GitHub | noreply@github.com | 2025-10-08T20:21:46+02:00 | | Disable CUDA host buffers on integrated GPUs (#16308) |
| 374 | d2ee056e1df0fdf646270fb5621e9f92084b59a7 | b2c08c9ec4cd59ed88c1ebdd94f109af5ccc978e | issixx | 46835150+issixx@users.noreply.github.com | 2025-10-08T17:20:18+09:00 | GitHub | noreply@github.com | 2025-10-08T11:20:18+03:00 | | server : fix cancel pending task (#16467) |
| 375 | b2c08c9ec4cd59ed88c1ebdd94f109af5ccc978e | 7fdd16b432d247121fe5fe7b21f8805f85266c85 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-08T10:57:53+03:00 | GitHub | noreply@github.com | 2025-10-08T10:57:53+03:00 | | metal : mark FA blocks (#16372) |
| 376 | 7fdd16b432d247121fe5fe7b21f8805f85266c85 | 74b8fc17f92ada295a648e3c5eb28f46bca7d892 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-08T10:57:29+03:00 | GitHub | noreply@github.com | 2025-10-08T10:57:29+03:00 | | server : improve context checkpoint logic (#16440) |
| 377 | 74b8fc17f92ada295a648e3c5eb28f46bca7d892 | aeaf8a36f06b5810f5ae4bbefe26edb33925cf5e | Reese Levine | reeselevine1@gmail.com | 2025-10-07T13:48:56-07:00 | GitHub | noreply@github.com | 2025-10-07T13:48:56-07:00 | | ggml webgpu: profiling, CI updates, reworking of command submission (#16452) |
| 378 | aeaf8a36f06b5810f5ae4bbefe26edb33925cf5e | df1b612e29ba97a2e67db339b1e8c7465702b7e8 | Tarek Dakhran | tarek@liquid.ai | 2025-10-07T20:03:35+02:00 | GitHub | noreply@github.com | 2025-10-07T20:03:35+02:00 | | llama : support LiquidAI LFM2-MoE hybrid model (#16464) |
| 379 | df1b612e29ba97a2e67db339b1e8c7465702b7e8 | 4e0388aa8a1f4ef1065701dc9c3947aea0a1b9a5 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T15:57:14+03:00 | GitHub | noreply@github.com | 2025-10-07T15:57:14+03:00 | | server : add `/v1/health` endpoint (#16461) |
| 380 | 4e0388aa8a1f4ef1065701dc9c3947aea0a1b9a5 | ef4c5b87ea2556ff8ca99cca3abdf48bdbca22f2 | Sascha Rogmann | 59577610+srogmann@users.noreply.github.com | 2025-10-07T11:11:08+02:00 | GitHub | noreply@github.com | 2025-10-07T11:11:08+02:00 | | webui : added download action (#13552) (#16282) |
| 381 | ef4c5b87ea2556ff8ca99cca3abdf48bdbca22f2 | c61ae20d05bd4fdd8551311325a2845336449426 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T10:32:32+03:00 | GitHub | noreply@github.com | 2025-10-07T10:32:32+03:00 | | presets : fix pooling param for embedding models (#16455) |
| 382 | c61ae20d05bd4fdd8551311325a2845336449426 | 0123ff38f53d34752f29239a29d0e40a6dc4110f | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-07T09:59:13+03:00 | GitHub | noreply@github.com | 2025-10-07T06:59:13Z | | rpc : update documentation (#16441) |
| 383 | 0123ff38f53d34752f29239a29d0e40a6dc4110f | 0a319bb75ed29d968e2a9b544011b09ccb932915 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T08:24:17+03:00 | GitHub | noreply@github.com | 2025-10-07T08:24:17+03:00 | | memory : use sequential equal splits for recurrent modules (#16442) |
| 384 | 0a319bb75ed29d968e2a9b544011b09ccb932915 | 1d6092fc72f4d10f4486ac95edfd414bc08b62b8 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T08:23:30+03:00 | GitHub | noreply@github.com | 2025-10-07T08:23:30+03:00 | | metal : add support for non-padded FA KV (#16148) |
| 385 | 1d6092fc72f4d10f4486ac95edfd414bc08b62b8 | 8ae32dc9ecd7aeebf3a5b43557e0552a0a04cd4f | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T08:22:35+03:00 | GitHub | noreply@github.com | 2025-10-07T08:22:35+03:00 | | tests : add -INF blocks to the KQ mask in the FA tests (#16380) |
| 386 | 8ae32dc9ecd7aeebf3a5b43557e0552a0a04cd4f | 3df2244df40c67dfd6ad548b40ccc507a066af2b | Georgi Gerganov | ggerganov@gmail.com | 2025-10-07T08:21:40+03:00 | GitHub | noreply@github.com | 2025-10-07T08:21:40+03:00 | | metal : various optimizations + refactoring (#16446) |
| 387 | 3df2244df40c67dfd6ad548b40ccc507a066af2b | c08002a1988348403a5fd59b1fa3de3a10a6f92f | Gadflyii | 34758915+Gadflyii@users.noreply.github.com | 2025-10-06T12:55:53-05:00 | GitHub | noreply@github.com | 2025-10-06T19:55:53+02:00 | | llama : add --no-host to disable host buffers (#16310) |
| 388 | c08002a1988348403a5fd59b1fa3de3a10a6f92f | 3a002afafa8e06555f11aed88b02d055d8a166f3 | Gabe Goodhart | ghart@us.ibm.com | 2025-10-06T10:59:40-06:00 | GitHub | noreply@github.com | 2025-10-06T18:59:40+02:00 | | chat : Granite Docling stopping (#16438) |
| 389 | 3a002afafa8e06555f11aed88b02d055d8a166f3 | a23b9bdbd3b64ce172f9962249f432d01aea7437 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-06T17:40:21+02:00 | GitHub | noreply@github.com | 2025-10-06T17:40:21+02:00 | | ci : refactor sdk caching to minimize storage (#16414) |
| 390 | a23b9bdbd3b64ce172f9962249f432d01aea7437 | 04e632a4aab8e6bfff0f8bc216b36ceb1e199ff9 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-06T16:05:27+03:00 | GitHub | noreply@github.com | 2025-10-06T16:05:27+03:00 | | ggml : fix unaligned access in AMX code (#16315) |
| 391 | 04e632a4aab8e6bfff0f8bc216b36ceb1e199ff9 | a80ff183abe4e5a76316257ffa597da41b3b6fa0 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-06T14:56:59+02:00 | GitHub | noreply@github.com | 2025-10-06T14:56:59+02:00 | | ci : remove missing reranker model files (#16444) |
| 392 | a80ff183abe4e5a76316257ffa597da41b3b6fa0 | 1d49ca37594fb49db6aa9518ba7c512e5ccd0108 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-06T14:17:12+02:00 | GitHub | noreply@github.com | 2025-10-06T14:17:12+02:00 | | ggml-cpu : fix leftover handling in ggml_vec_scale_f32 for SVE (#16443) |
| 393 | 1d49ca37594fb49db6aa9518ba7c512e5ccd0108 | c5fef0fcea3b978a2318b1af170209ecec7c37b4 | Yuannan | yuannan@users.noreply.github.com | 2025-10-06T09:29:56Z | GitHub | noreply@github.com | 2025-10-06T12:29:56+03:00 | | nix : removed metal for nix (#16118) |
| 394 | c5fef0fcea3b978a2318b1af170209ecec7c37b4 | ca71fb9b368e3db96e028f80c4c9df6b6b370edd | Oleksandr Kuvshynov | 661042+okuvshynov@users.noreply.github.com | 2025-10-06T03:53:31-04:00 | GitHub | noreply@github.com | 2025-10-06T10:53:31+03:00 | | server: update readme to mention n_past_max metric (#16436) |
| 395 | ca71fb9b368e3db96e028f80c4c9df6b6b370edd | 35266573b968e1c947b367782fb4b3eddbb4f3c0 | Gabe Goodhart | ghart@us.ibm.com | 2025-10-05T06:57:47-06:00 | GitHub | noreply@github.com | 2025-10-05T14:57:47+02:00 | | model : Granite docling + Idefics3 preprocessing (SmolVLM) (#16206) |
| 396 | 35266573b968e1c947b367782fb4b3eddbb4f3c0 | 86df2c9ae4f2f1ee63d2558a9dc797b98524639b | Reese Levine | reeselevine1@gmail.com | 2025-10-04T20:59:31-07:00 | GitHub | noreply@github.com | 2025-10-04T20:59:31-07:00 | | ggml webgpu: actually add softmax, fix rms_norm offset (#16400) |
| 397 | 86df2c9ae4f2f1ee63d2558a9dc797b98524639b | f39283960b58a92ecc0c72567711318b20e22b55 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-10-04T20:04:27Z | GitHub | noreply@github.com | 2025-10-04T22:04:27+02:00 | | vulkan: use a more appropriate amount of threads when generating shaders (#16418) |
| 398 | f39283960b58a92ecc0c72567711318b20e22b55 | 898acba6816ad23b6a9491347d30e7570bffadfd | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-04T16:22:45+03:00 | GitHub | noreply@github.com | 2025-10-04T16:22:45+03:00 | | rpc : check src buffer when copying tensor (#16421) |
| 399 | 898acba6816ad23b6a9491347d30e7570bffadfd | e29acf74fea996014380d59d31aa504ae8964258 | Radoslav Gerganov | rgerganov@gmail.com | 2025-10-04T12:49:16+03:00 | GitHub | noreply@github.com | 2025-10-04T12:49:16+03:00 | | rpc : add support for multiple devices (#16276) |
| 400 | e29acf74fea996014380d59d31aa504ae8964258 | 128d522c04286e019666bd6ee4d18e3fbf8772e2 | Acly | aclysia@gmail.com | 2025-10-04T11:42:56+02:00 | GitHub | noreply@github.com | 2025-10-04T11:42:56+02:00 | | vulkan : incremental shader builds (#16341) |
| 401 | 128d522c04286e019666bd6ee4d18e3fbf8772e2 | f6dcda390004b627ef30af378d0c01ad2519289e | Pascal | admin@serveurperso.com | 2025-10-03T20:51:48+02:00 | GitHub | noreply@github.com | 2025-10-03T21:51:48+03:00 | | chat : support Magistral thinking (#16413) |
| 402 | f6dcda390004b627ef30af378d0c01ad2519289e | 606a73f53175077429484b23dcf799f69a31d0bd | ddh0 | dylanhalladay02@icloud.com | 2025-10-03T13:34:51-05:00 | GitHub | noreply@github.com | 2025-10-03T21:34:51+03:00 | | server : context checkpointing for hybrid and recurrent models (#16382) |
| 403 | 606a73f53175077429484b23dcf799f69a31d0bd | 946f71ed9ade07e319859b5ce656144140e066fb | Georgi Gerganov | ggerganov@gmail.com | 2025-10-03T19:18:56+03:00 | GitHub | noreply@github.com | 2025-10-03T19:18:56+03:00 | | metal : fix loop bound in ggml_mem_ranges (#16412) |
| 404 | 946f71ed9ade07e319859b5ce656144140e066fb | 638d330246954e88dffc22ce01fec15e6894e544 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-03T14:40:25+02:00 | GitHub | noreply@github.com | 2025-10-03T14:40:25+02:00 | | llama : fix shapes for bert/mpt q/k norm (#16409) |
| 405 | 638d330246954e88dffc22ce01fec15e6894e544 | 84c8e305e8010a1a3d43bdd0a25f737ac67809a4 | Acly | aclysia@gmail.com | 2025-10-03T13:49:08+02:00 | GitHub | noreply@github.com | 2025-10-03T13:49:08+02:00 | | ggml : fix graph reallocation with multiple chunks (#16396) |
| 406 | 84c8e305e8010a1a3d43bdd0a25f737ac67809a4 | 2aaf0a2a2056d75d0dd53ab8a181473760e6ab22 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-03T12:51:40+02:00 | GitHub | noreply@github.com | 2025-10-03T12:51:40+02:00 | | Fix missing messages on sibling navigation (#16408) |
| 407 | 2aaf0a2a2056d75d0dd53ab8a181473760e6ab22 | 0e1f83855609d73beaf05d818640b6cfd39d287b | Jeff Bolz | jbolz@nvidia.com | 2025-10-03T05:50:46-05:00 | GitHub | noreply@github.com | 2025-10-03T12:50:46+02:00 | | vulkan: Replace uses of maxMemoryAllocationSize and VK_WHOLE_SIZE (#16354) |
| 408 | 0e1f83855609d73beaf05d818640b6cfd39d287b | ad126479c25cf983a0f994a08ba0911cf49ed62b | Jeff Bolz | jbolz@nvidia.com | 2025-10-03T04:52:46-05:00 | GitHub | noreply@github.com | 2025-10-03T11:52:46+02:00 | | vulkan: Fix FA coopmat1 invalid array indexing (#16365) |
| 409 | ad126479c25cf983a0f994a08ba0911cf49ed62b | 77233277c912d495d1069543ca540c21a2f9e5b4 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-10-03T11:45:16+02:00 | GitHub | noreply@github.com | 2025-10-03T11:45:16+02:00 | | ci : change macos-13 to macos-15-intel (#16401) |
| 410 | 77233277c912d495d1069543ca540c21a2f9e5b4 | e308efda8e54e3f07578937a45a5d83624eb7970 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-03T11:30:39+02:00 | GitHub | noreply@github.com | 2025-10-03T11:30:39+02:00 | | Capture model name only after first token (streaming) or completed request (#16405) |
| 411 | e308efda8e54e3f07578937a45a5d83624eb7970 | 136bda78c5679b266362b39c58338fff46fd2592 | Jeff Bolz | jbolz@nvidia.com | 2025-10-03T03:33:08-05:00 | GitHub | noreply@github.com | 2025-10-03T10:33:08+02:00 | | vulkan: in flash attention, bounds check against nem1 (don't rely on GGML_KQ_MASK_PAD) (#16316) |
| 412 | 136bda78c5679b266362b39c58338fff46fd2592 | 5113efd34ceda709292de26c72716c49e024fb32 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-03T09:11:34+02:00 | GitHub | noreply@github.com | 2025-10-03T10:11:34+03:00 | | webui : Fix messages payload sent to chat completions (#16402) |
| 413 | 5113efd34ceda709292de26c72716c49e024fb32 | d64c8104f090b27b1f99e8da5995ffcfa6b726e2 | Pascal | admin@serveurperso.com | 2025-10-03T08:01:31+02:00 | GitHub | noreply@github.com | 2025-10-03T08:01:31+02:00 | | fix: track viewportHeight via window.innerHeight to avoid unwanted scrolling (#16356) |
| 414 | d64c8104f090b27b1f99e8da5995ffcfa6b726e2 | ef07a4090672a3438d7f64f197795d7dc1c18957 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-02T20:10:12+02:00 | GitHub | noreply@github.com | 2025-10-02T20:10:12+02:00 | | test-barrier : do not use more threads than physically available (#16389) |
| 415 | ef07a4090672a3438d7f64f197795d7dc1c18957 | 34fcc5a4ace8c69476ef2ea3857f39a60334acc4 | Reese Levine | reeselevine1@gmail.com | 2025-10-02T11:00:31-07:00 | GitHub | noreply@github.com | 2025-10-02T11:00:31-07:00 | | ggml webgpu: add support for soft_max, optimize rms_norm (#16357) |
| 416 | 34fcc5a4ace8c69476ef2ea3857f39a60334acc4 | 91a2a5655658bb9ab77894716b82fae7ecb4b4d1 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-10-02T19:43:22+02:00 | GitHub | noreply@github.com | 2025-10-02T20:43:22+03:00 | | model : Apertus model implementation (#15852) |
| 417 | 91a2a5655658bb9ab77894716b82fae7ecb4b4d1 | 72ee736c447be4ededb51ab97904b4d33c90c261 | R0CKSTAR | yeahdongcn@gmail.com | 2025-10-02T21:29:56+08:00 | GitHub | noreply@github.com | 2025-10-02T16:29:56+03:00 | | musa: update compile flags (#16265) |
| 418 | 72ee736c447be4ededb51ab97904b4d33c90c261 | f09aefaa84d7f4d5df3f400f67944b94fef5b795 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-02T13:51:36+02:00 | GitHub | noreply@github.com | 2025-10-02T13:51:36+02:00 | | ci : fix ubuntu-latest-cmake-rpc (disable ccache) (#16388) |
| 419 | f09aefaa84d7f4d5df3f400f67944b94fef5b795 | bbd32bc0384fbfcf07369617de58856b1e0e95a3 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-10-02T08:10:07Z | GitHub | noreply@github.com | 2025-10-02T10:10:07+02:00 | | ci: update vulkan ci (#16294) |
| 420 | bbd32bc0384fbfcf07369617de58856b1e0e95a3 | 2be72c2b121ee99f33927149265ce6073ade9e59 | Georgi Gerganov | ggerganov@gmail.com | 2025-10-02T10:35:43+03:00 | GitHub | noreply@github.com | 2025-10-02T10:35:43+03:00 | | ci : fix clean-up of old logs (#16381) |
| 421 | 2be72c2b121ee99f33927149265ce6073ade9e59 | 95ce0985449c52fb9d6849b95563f7cf933ea0e3 | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-10-02T15:16:25+08:00 | GitHub | noreply@github.com | 2025-10-02T10:16:25+03:00 | | SYCL: Update to oneAPI 2025.2 (#16371) |
| 422 | 95ce0985449c52fb9d6849b95563f7cf933ea0e3 | c8dedc9999eccf7821a9fe5b29f10e8d075e2217 | uvos | carl@uvos.xyz | 2025-10-02T05:52:59+02:00 | GitHub | noreply@github.com | 2025-10-02T05:52:59+02:00 | | HIP: add IMbackK to codeowner (#16375) |
| 423 | c8dedc9999eccf7821a9fe5b29f10e8d075e2217 | e95fec640f43623911a2cd5bda8b19b1898c530c | uvos | carl@uvos.xyz | 2025-10-01T23:32:39+02:00 | GitHub | noreply@github.com | 2025-10-01T23:32:39+02:00 | | CI: reenable cdna in rocm docker builds (#16376) |
| 424 | e95fec640f43623911a2cd5bda8b19b1898c530c | ded67b94446ef4f7fd988dbde7a12deef9870c13 | uvos | carl@uvos.xyz | 2025-10-01T23:09:25+02:00 | GitHub | noreply@github.com | 2025-10-01T23:09:25+02:00 | | HIP: Disable ROCWMMA fattn on CDNA when compiled against ROCWMMA 2.0.0 (#16221) |
| 425 | ded67b94446ef4f7fd988dbde7a12deef9870c13 | 1fe4e38cc20af058ed320bd46cac934991190056 | Shunta Saito | shunta.saito@gmail.com | 2025-10-02T06:08:15+09:00 | GitHub | noreply@github.com | 2025-10-01T23:08:15+02:00 | | llama : parameter conversion and loading fixes for PLaMo2 variants (#16075) |
| 426 | 1fe4e38cc20af058ed320bd46cac934991190056 | 4201deae9c2ae4732db8957b6ce0808d02ec597c | uvos | carl@uvos.xyz | 2025-10-01T20:18:03+02:00 | GitHub | noreply@github.com | 2025-10-01T20:18:03+02:00 | | ci: Properly install rocwmma for hip builds (#16305) |
| 427 | 4201deae9c2ae4732db8957b6ce0808d02ec597c | 764799279f801c696503bacca490b4358587d94c | Adrien Gallouët | adrien@gallouet.fr | 2025-10-01T19:22:18+02:00 | GitHub | noreply@github.com | 2025-10-01T20:22:18+03:00 | | common: introduce http.h for httplib-based client (#16373) |
| 428 | 764799279f801c696503bacca490b4358587d94c | 2a9b63383a448ec18c754dfdc6e95cb853940a52 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-01T18:18:10+02:00 | GitHub | noreply@github.com | 2025-10-01T18:18:10+02:00 | | Conversation action dialogs as singletons from Chat Sidebar + apply conditional rendering for Actions Dropdown for Chat Conversation Items (#16369) |
| 429 | 2a9b63383a448ec18c754dfdc6e95cb853940a52 | 1104ca1a1c926b832274694a2773a17066982ad0 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-01T15:54:42+02:00 | GitHub | noreply@github.com | 2025-10-01T15:54:42+02:00 | | Improve code block color theming (#16325) |
| 430 | 1104ca1a1c926b832274694a2773a17066982ad0 | 4f1575921cac9b489fe8f8bbf20aff985e7da45e | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-10-01T14:09:52+02:00 | GitHub | noreply@github.com | 2025-10-01T14:09:52+02:00 | | ci : use registry cache for docker builds (#16366) |
| 431 | 132d673554e65282e92505f4f77401ab971528bb | aa9538a63aef35f0224b2502f13b19d650a8ce04 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-10-01T07:56:36Z | GitHub | noreply@github.com | 2025-10-01T09:56:36+02:00 | | vulkan: make ggml_vk_default_dispatcher support older vulkan headers (#16345) |
| 432 | aa9538a63aef35f0224b2502f13b19d650a8ce04 | e74c92e84236b2bab3f3c77bee4ead94928be360 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-10-01T07:40:26+02:00 | GitHub | noreply@github.com | 2025-10-01T08:40:26+03:00 | | webui: Remove running `llama-server` within WebUI `dev.sh` script (#16363) |
| 433 | e74c92e84236b2bab3f3c77bee4ead94928be360 | b2ba81dbe07b6dbea9c96b13346c66973dede32c | Bartowski | 3266127+bartowski1182@users.noreply.github.com | 2025-09-30T16:24:36-04:00 | GitHub | noreply@github.com | 2025-09-30T22:24:36+02:00 | | model : support GLM 4.6 (make a few NextN/MTP tensors not required) (#16359) |
| 434 | b2ba81dbe07b6dbea9c96b13346c66973dede32c | bf6f3b3a1965d70e07ca94aab7b01268fe483e96 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-30T21:41:42+02:00 | GitHub | noreply@github.com | 2025-09-30T21:41:42+02:00 | | ci : fix ccache key for ubuntu-cpu-cmake (#16355) |
| 435 | bf6f3b3a1965d70e07ca94aab7b01268fe483e96 | 7c156df4148eb524b22f7f363bfd2df00f264e0b | Adrien Gallouët | angt@huggingface.co | 2025-09-30T19:52:41+02:00 | GitHub | noreply@github.com | 2025-09-30T20:52:41+03:00 | | common : disable progress bar without a tty (#16352) |
| 436 | 7c156df4148eb524b22f7f363bfd2df00f264e0b | 16b0ca0d2e63fc1a0a43795f8552f9cb61d9f7a5 | lhez | lih@qti.qualcomm.com | 2025-09-30T10:45:45-07:00 | GitHub | noreply@github.com | 2025-09-30T10:45:45-07:00 | | opencl: support pad_ext (#15888) |
| 437 | 16b0ca0d2e63fc1a0a43795f8552f9cb61d9f7a5 | 8d78cd2613ccdeb3cc86f59bc8f9ddd31cfbd3ed | Pascal | admin@serveurperso.com | 2025-09-30T19:18:54+02:00 | GitHub | noreply@github.com | 2025-09-30T19:18:54+02:00 | | Chatapi ignore empty sampling (#16330) |
| 438 | 8d78cd2613ccdeb3cc86f59bc8f9ddd31cfbd3ed | d1c84a662daa91be975863913a59975b11458141 | Reese Levine | reeselevine1@gmail.com | 2025-09-30T09:57:51-07:00 | GitHub | noreply@github.com | 2025-09-30T09:57:51-07:00 | | ggml webgpu: support for rope,div,sub,glu,scale,cont operators (#16187) |
| 439 | d1c84a662daa91be975863913a59975b11458141 | 364a7a6d4a786e98947c8a90430ea581213c0ba9 | lhez | lih@qti.qualcomm.com | 2025-09-30T09:55:13-07:00 | GitHub | noreply@github.com | 2025-09-30T09:55:13-07:00 | | opencl: support ne3 in get_rows (#15866) |
| 440 | 364a7a6d4a786e98947c8a90430ea581213c0ba9 | 2df5bcf357dba0c49b4df1d684b2ffc9e88b7054 | Adrien Gallouët | angt@huggingface.co | 2025-09-30T16:39:44+02:00 | GitHub | noreply@github.com | 2025-09-30T17:39:44+03:00 | | common : remove common_has_curl() (#16351) |
| 441 | 2df5bcf357dba0c49b4df1d684b2ffc9e88b7054 | 075c01567bb7e93965ae60b7e119182c3e68f3e9 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-30T15:38:01+02:00 | GitHub | noreply@github.com | 2025-09-30T15:38:01+02:00 | | ci : disable ccache for android (#16348) |
| 442 | 075c01567bb7e93965ae60b7e119182c3e68f3e9 | a014310374a16f9204f2bcc1b458fc1eda67e469 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-30T13:42:39+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-30T13:53:55+03:00 | | ggml : bump version to 0.9.4 (ggml/1363) |
| 443 | 98c8269abefdfa8fde47dfb2769ceebc8b704e86 | 2790ecc304e384d9075c55d0628d8bf440025dc4 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-30T17:42:57+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-30T17:44:27+08:00 | | server: adapt to calrt |
| 444 | 2790ecc304e384d9075c55d0628d8bf440025dc4 | f528e97d894b4d93934a0711a00253eea980e799 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-30T17:37:58+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-30T17:44:27+08:00 | | adapt to calrt |
| 445 | f528e97d894b4d93934a0711a00253eea980e799 | dcbe998a6d8f13b6c9cb36bbe7b355ced2ac3842 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-11T14:13:35+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-30T17:44:27+08:00 | | add calrt mtmd support |
| 446 | a014310374a16f9204f2bcc1b458fc1eda67e469 | 35fb82497ec6c5904b0adb7e1c881a76c1c692db | anavp-nvidia | anavp@nvidia.com | 2025-09-30T08:13:22Z | GitHub | noreply@github.com | 2025-09-30T11:13:22+03:00 | | cuda : Enable CUDA Graph usage for Nemotron Nano v2 (NemotronH) (#16328) |
| 447 | 35fb82497ec6c5904b0adb7e1c881a76c1c692db | 3c62aed89fbe3e15f6f82035e805a0a1add26306 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-30T11:03:23+03:00 | GitHub | noreply@github.com | 2025-09-30T11:03:23+03:00 | | metal : dynamic simdgroups for MV kernels (#16340) |
| 448 | 3c62aed89fbe3e15f6f82035e805a0a1add26306 | f1eb1cb1eba042b2d583234e80f65e2e225d8995 | Adrien Gallouët | angt@huggingface.co | 2025-09-30T09:36:33+02:00 | GitHub | noreply@github.com | 2025-09-30T10:36:33+03:00 | | common : simplify etag tracking by removing json (#16342) |
| 449 | f1eb1cb1eba042b2d583234e80f65e2e225d8995 | de41f2b7bffab7582944b213ce12b605f8cf6e0f | Charles Xu | charles.xu@arm.com | 2025-09-30T09:07:20+02:00 | GitHub | noreply@github.com | 2025-09-30T10:07:20+03:00 | | kleidiai : fix work size and threads sync for fp16 (#16246) |
| 450 | de41f2b7bffab7582944b213ce12b605f8cf6e0f | a74a0d69f34f52fa10d4f0a7ce749fb3490d0774 | lhez | lih@qti.qualcomm.com | 2025-09-29T22:30:16-07:00 | GitHub | noreply@github.com | 2025-09-30T08:30:16+03:00 | | codeowners: add codeowners for opencl backend (#16344) |
| 451 | a74a0d69f34f52fa10d4f0a7ce749fb3490d0774 | 5f7e166cbf7b9ca928c7fad990098ef32358ac75 | Jeff Bolz | jbolz@nvidia.com | 2025-09-29T19:26:34-05:00 | GitHub | noreply@github.com | 2025-09-29T19:26:34-05:00 | | tests: override test_set_rows::max_nmse_err to allow for occasional rounding differences (#16295) |
| 452 | 5f7e166cbf7b9ca928c7fad990098ef32358ac75 | d72f5f7ba260b546190338b0b76f2f152581424f | Pascal | admin@serveurperso.com | 2025-09-29T18:49:47+02:00 | GitHub | noreply@github.com | 2025-09-29T18:49:47+02:00 | | Fix thinking blocks with quotes + add handling `[THINK]...[/THINK]` blocks (#16326) |
| 453 | d72f5f7ba260b546190338b0b76f2f152581424f | b77e6c18e1a6fac5705ed95f03af5436d67484c1 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:51:48+03:00 | GitHub | noreply@github.com | 2025-09-29T17:51:48+03:00 | | ci : add AMD runners and workflows (#16249) |
| 454 | b77e6c18e1a6fac5705ed95f03af5436d67484c1 | 2ddd3f2356c313385fafc19a9f0ce678c8fe03ee | alex-spacemit | jinghui.huang@spacemit.com | 2025-09-29T22:50:44+08:00 | GitHub | noreply@github.com | 2025-09-29T17:50:44+03:00 | | ggml: riscv: add riscv spacemit backend (#15288) |
| 455 | 2ddd3f2356c313385fafc19a9f0ce678c8fe03ee | 4d3d455d3c8719eb8f206de328d1b9ba191efd7c | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T16:50:52+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | sync : ggml |
| 456 | 4d3d455d3c8719eb8f206de328d1b9ba191efd7c | c9b1c06467aca8d30fafcab51292bcb285df5b0d | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T16:49:11+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | sync : whisper.cpp (ggml/1359) |
| 457 | c9b1c06467aca8d30fafcab51292bcb285df5b0d | b6ae75afb49b23db3dfbc2b01e8aabfb047e6880 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-26T17:34:42+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | ggml : remove -dev suffix from release version (ggml/1355) |
| 458 | b6ae75afb49b23db3dfbc2b01e8aabfb047e6880 | b6dff20e2fe352093039e2359370c6bb2b505e0b | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-25T14:39:05+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | ggml : bump version to 0.9.3 (ggml/1353) |
| 459 | b6dff20e2fe352093039e2359370c6bb2b505e0b | 2db78c75e4ef7497313fda6dbbb3fcaf9ca0752d | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T16:44:23+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | ggml : prepare for development of 0.9.2-dev |
| 460 | 2db78c75e4ef7497313fda6dbbb3fcaf9ca0752d | 02463ab27b1379ff5e3d936ce8b3bfd356872ea6 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T16:44:23+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T17:43:58+03:00 | | ggml : bump version to 0.9.1 |
| 461 | 02463ab27b1379ff5e3d936ce8b3bfd356872ea6 | adc76347d73b3d915e946efa5de8d0ad9f3904c2 | Rafal Lewczuk | rafal.lewczuk@gmail.com | 2025-09-29T13:17:09+02:00 | GitHub | noreply@github.com | 2025-09-29T13:17:09+02:00 | | ggml-backend : add root cause in error message if loading backend library fails (#16172) |
| 462 | adc76347d73b3d915e946efa5de8d0ad9f3904c2 | 3a2bdcda0b31e51db954551e0cd49424133a84f3 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-29T11:09:00+02:00 | GitHub | noreply@github.com | 2025-09-29T11:09:00+02:00 | | ggml : check cuda and metal argsort limits and add test (#16323) |
| 463 | 3a2bdcda0b31e51db954551e0cd49424133a84f3 | 66bb7985c3fcefda7a74aae9f8321f70eb1646c4 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-29T10:37:20+02:00 | GitHub | noreply@github.com | 2025-09-29T10:37:20+02:00 | | Improve Mobile UI for dialogs and action dropdowns (#16222) |
| 464 | 66bb7985c3fcefda7a74aae9f8321f70eb1646c4 | 2f61c0f5bf8a620ca4c3872408803ab38cfb9613 | Pascal | admin@serveurperso.com | 2025-09-29T09:08:41+02:00 | GitHub | noreply@github.com | 2025-09-29T09:08:41+02:00 | | fix: preserved zero values in chat settings inputs and textareas by switching to nullish coalescing for field values and default placeholders (#16312) |
| 465 | 2f61c0f5bf8a620ca4c3872408803ab38cfb9613 | 3ffd0fae473c954bb3e67526b31262048fb508d4 | Vinkal | vinkal-chudgar@users.noreply.github.com | 2025-09-29T12:33:12+05:30 | GitHub | noreply@github.com | 2025-09-29T10:03:12+03:00 | | llama-cli: prevent spurious assistant token (#16202) |
| 466 | 3ffd0fae473c954bb3e67526b31262048fb508d4 | a4a0aa5ea2a88e4858996199fefd5439f17b481c | ddh0 | chemist-mulches-39@icloud.com | 2025-09-29T01:30:45-05:00 | GitHub | noreply@github.com | 2025-09-29T09:30:45+03:00 | | perplexity : show more kl-divergence data (#16321) |
| 467 | a4a0aa5ea2a88e4858996199fefd5439f17b481c | 92cd103f627015a95f2107f1c27dc218c7b8ec64 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-29T08:41:28+03:00 | GitHub | noreply@github.com | 2025-09-29T08:41:28+03:00 | | ggml : fix dependencies for ggml_set_rows (#16318) |
| 468 | 92cd103f627015a95f2107f1c27dc218c7b8ec64 | b887d2f3413ac231e3cb5925260c39902af4a70c | Jeff Bolz | jbolz@nvidia.com | 2025-09-28T23:50:37-05:00 | GitHub | noreply@github.com | 2025-09-29T06:50:37+02:00 | | vulkan: Fix validation failure in quantized flash attention (#16292) |
| 469 | b887d2f3413ac231e3cb5925260c39902af4a70c | bd0af02fc96c2057726f33c0f0daf7bb8f3e462a | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-28T23:15:03+02:00 | GitHub | noreply@github.com | 2025-09-28T23:15:03+02:00 | | ggml : fix GGML_F32_VEC_FMA argument order in ggml_vec_mad1_f32 (#16307) |
| 470 | bd0af02fc96c2057726f33c0f0daf7bb8f3e462a | d9e0e7c8194dfd7d23bf3a86608c9ece68d77c93 | crat0z | 11581854+crat0z@users.noreply.github.com | 2025-09-28T14:13:50-04:00 | GitHub | noreply@github.com | 2025-09-28T21:13:50+03:00 | | common : fix reasoning before forced tool call via tool_choice = required (#16264) |
| 471 | d9e0e7c8194dfd7d23bf3a86608c9ece68d77c93 | 0124ac989f7e7bf08803788f66dbe4106bdcdd58 | R0CKSTAR | yeahdongcn@gmail.com | 2025-09-28T22:38:15+08:00 | GitHub | noreply@github.com | 2025-09-28T16:38:15+02:00 | | ci : fix musa docker build (#16306) |
| 472 | 0124ac989f7e7bf08803788f66dbe4106bdcdd58 | 2811c65286ae954bec87049f75b86dc022006dcc | Aaron Teo | aaron.teo1@ibm.com | 2025-09-28T19:25:58+08:00 | GitHub | noreply@github.com | 2025-09-28T19:25:58+08:00 | | devops: switch to using ubuntu-22.04-s390x image (#16302) |
| 473 | 2811c65286ae954bec87049f75b86dc022006dcc | d8359f5fde480da030bf75c7711573c7c4d993ba | Imad Saddik | 79410781+ImadSaddik@users.noreply.github.com | 2025-09-28T12:04:46+01:00 | GitHub | noreply@github.com | 2025-09-28T13:04:46+02:00 | | Fixed a few typos in the README of the LLaMA.cpp HTTP Server [no ci] (#16297) |
| 474 | d8359f5fde480da030bf75c7711573c7c4d993ba | 6a2c6145a0b91b40eb3c3dba7b20ccc4b270490f | Jeff Bolz | jbolz@nvidia.com | 2025-09-28T01:38:37-05:00 | GitHub | noreply@github.com | 2025-09-28T08:38:37+02:00 | | vulkan: 64-bit im2col (#16135) |
| 475 | 6a2c6145a0b91b40eb3c3dba7b20ccc4b270490f | 3b53634fe35771e2e318227aa81585726bae7234 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-28T09:34:44+03:00 | GitHub | noreply@github.com | 2025-09-28T09:34:44+03:00 | | metal : extend mat-mat multiplication support (#16225) |
| 476 | 3b53634fe35771e2e318227aa81585726bae7234 | 1384abf8b8d5894d32fada453ccf4d196ffba7de | Georgi Gerganov | ggerganov@gmail.com | 2025-09-28T09:34:05+03:00 | GitHub | noreply@github.com | 2025-09-28T09:34:05+03:00 | | metal : fuse non-sequential nodes (#16102) |
| 477 | 1384abf8b8d5894d32fada453ccf4d196ffba7de | e6d65fb02d553bd79cad94e517cdca18b687788d | Jeff Bolz | jbolz@nvidia.com | 2025-09-27T20:36:34-05:00 | GitHub | noreply@github.com | 2025-09-27T20:36:34-05:00 | | vulkan: handle mat_mul with A matrix > 4GB (#16176) |
| 478 | e6d65fb02d553bd79cad94e517cdca18b687788d | 8656f5de688cddcaea1d6174535eb60ee23ef6a0 | Jeff Bolz | jbolz@nvidia.com | 2025-09-27T16:43:39-04:00 | GitHub | noreply@github.com | 2025-09-27T22:43:39+02:00 | | vulkan: support arbitrary KV dimension in flash attention (#16160) |
| 479 | 8656f5de688cddcaea1d6174535eb60ee23ef6a0 | 4807e8f96a61b2adccebd5e57444c94d18de7264 | Acly | aclysia@gmail.com | 2025-09-27T22:41:03+02:00 | GitHub | noreply@github.com | 2025-09-27T22:41:03+02:00 | | vulkan : make the vulkan.hpp dynamic dispatcher instance private (#16224) |
| 480 | 4807e8f96a61b2adccebd5e57444c94d18de7264 | c0bfc57af421f8fd63c946c13b7666aed82560e2 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-27T19:56:40+02:00 | GitHub | noreply@github.com | 2025-09-27T19:56:40+02:00 | | Show message actions by default (#16289) |
| 481 | c0bfc57af421f8fd63c946c13b7666aed82560e2 | 75a3a6c2cd0002ba40e2dcc92007bc9fdbc69f1a | Aman Gupta | amangupta052@gmail.com | 2025-09-28T00:49:32+08:00 | GitHub | noreply@github.com | 2025-09-27T18:49:32+02:00 | | CUDA: mul_mat_id for mmf for bs <= 64 for f16 and bs <= 32 for f32 (#16277) |
| 482 | 75a3a6c2cd0002ba40e2dcc92007bc9fdbc69f1a | 0499b29c6f64c705faaf5860dc4600fca23671f4 | Johannes Gäßler | johannesg@5d6.de | 2025-09-27T18:45:07+02:00 | GitHub | noreply@github.com | 2025-09-27T18:45:07+02:00 | | CUDA: refactor and deduplicate vector FA kernels (#16208) |
| 483 | 0499b29c6f64c705faaf5860dc4600fca23671f4 | 234e2ff8ed09716fb553437596779399bee31b11 | Dmytro Minochkin | dmytro.minochkin@gmail.com | 2025-09-27T19:26:46+03:00 | GitHub | noreply@github.com | 2025-09-27T18:26:46+02:00 | | vulkan: throw system error instead of SIGABRT during init on older devices (#16156) |
| 484 | 234e2ff8ed09716fb553437596779399bee31b11 | 3f81b4e91c1d5f098148af117e3f13cf4b077f52 | Adrien Gallouët | angt@huggingface.co | 2025-09-27T18:17:08+02:00 | GitHub | noreply@github.com | 2025-09-27T19:17:08+03:00 | | server : remove old LLAMA_SERVER_SSL (#16290) |
| 485 | 3f81b4e91c1d5f098148af117e3f13cf4b077f52 | ace6a54565444b6377bee8e7ac693238e7766279 | Jeff Bolz | jbolz@nvidia.com | 2025-09-27T06:36:11-04:00 | GitHub | noreply@github.com | 2025-09-27T12:36:11+02:00 | | vulkan: support GET_ROWS for k-quants (#16235) |
| 486 | ace6a54565444b6377bee8e7ac693238e7766279 | 72b24d96c6888c609d562779a23787304ae4609c | Adrien Gallouët | angt@huggingface.co | 2025-09-27T11:12:46+02:00 | GitHub | noreply@github.com | 2025-09-27T12:12:46+03:00 | | build : add LLAMA_OPENSSL option (#16287) |
| 487 | 72b24d96c6888c609d562779a23787304ae4609c | 624207e676ab5eb3ce7af631902bb45fb73a8359 | Vinkal | vinkal-chudgar@users.noreply.github.com | 2025-09-27T02:58:29+05:30 | GitHub | noreply@github.com | 2025-09-26T23:28:29+02:00 | | model : make minicpm embedding_scale, residual_scale and logit_scale optional with legacy defaults (#16273) |
| 488 | 624207e676ab5eb3ce7af631902bb45fb73a8359 | 807e8c6d310952f2f5656afc63dab9d7083dcb5c | Aaron Teo | aaron.teo1@ibm.com | 2025-09-27T02:03:33+08:00 | GitHub | noreply@github.com | 2025-09-27T02:03:33+08:00 | | devops: add s390x & ppc64le CI (#15925) |
| 489 | 807e8c6d310952f2f5656afc63dab9d7083dcb5c | 1a189278944d030211f336e103e96b65f976c361 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-26T19:25:29+02:00 | GitHub | noreply@github.com | 2025-09-26T19:25:29+02:00 | | Enhance text file detection logic for file attachments (#16199) |
| 490 | 1a189278944d030211f336e103e96b65f976c361 | e0539eb6aed346d4b25a6ea019044e88771e7690 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-26T18:35:42+02:00 | GitHub | noreply@github.com | 2025-09-26T18:35:42+02:00 | | Allow viewing conversations even when llama server is down (#16255) |
| 491 | e0539eb6aed346d4b25a6ea019044e88771e7690 | 5d0a40f390732cbf85d8a3b7b0fc3cbebffe780a | Isaac McFadyen | isaac@imcf.me | 2025-09-26T11:36:48-04:00 | GitHub | noreply@github.com | 2025-09-26T18:36:48+03:00 | | webui: switch to hash-based routing (alternative of #16079) (#16157) |
| 492 | 5d0a40f390732cbf85d8a3b7b0fc3cbebffe780a | d12a9836597b1ae4440d64d5a6ed727d27ac0702 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-26T15:59:07+02:00 | GitHub | noreply@github.com | 2025-09-26T15:59:07+02:00 | | Always show message actions for mobile UI + improvements for user message sizing (#16076) |
| 493 | d12a9836597b1ae4440d64d5a6ed727d27ac0702 | cc1cfa277b3aae1f6cc9180472072336597a78d4 | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-26T16:09:34+03:00 | GitHub | noreply@github.com | 2025-09-26T16:09:34+03:00 | | codeowners : add rgerganov as owner of RPC [no ci] (#16279) |
| 494 | cc1cfa277b3aae1f6cc9180472072336597a78d4 | 54dbc37053f7d75ea5b0631d5ee88f19391e4314 | Aleksei Nikiforov | 103434461+AlekseiNikiforovIBM@users.noreply.github.com | 2025-09-26T15:00:44+02:00 | GitHub | noreply@github.com | 2025-09-26T15:00:44+02:00 | | mtmd : fix uninitialized variable in bicubic_resize (#16275) |
| 495 | 54dbc37053f7d75ea5b0631d5ee88f19391e4314 | b995a10760cb93d23d617d76ecb82a5f95b5e0d3 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-26T14:14:28+03:00 | GitHub | noreply@github.com | 2025-09-26T14:14:28+03:00 | | metal : report OOM errors (#16274) |
| 496 | b995a10760cb93d23d617d76ecb82a5f95b5e0d3 | 4710dd31bbcef79d04f85a3a6a8c7d9439c5c79a | Adrien Gallouët | angt@huggingface.co | 2025-09-26T13:12:19+02:00 | GitHub | noreply@github.com | 2025-09-26T14:12:19+03:00 | | common : use cpp-httplib as a cURL alternative for downloads (#16185) |
| 497 | 4710dd31bbcef79d04f85a3a6a8c7d9439c5c79a | 9b26511857ac09ae69ab485168fe2d3ee5fb1d6e | Adrien Gallouët | angt@huggingface.co | 2025-09-26T12:39:35+02:00 | GitHub | noreply@github.com | 2025-09-26T13:39:35+03:00 | | build : fix build-ios-device (#16257) |
| 498 | 9b26511857ac09ae69ab485168fe2d3ee5fb1d6e | 00217cd41388328c93e9e9644921c4319bb03bcc | Aaron Teo | aaron.teo1@ibm.com | 2025-09-26T18:27:25+08:00 | GitHub | noreply@github.com | 2025-09-26T13:27:25+03:00 | | ggml-cpu: implement MXFP4 SIMD for s390x (#16193) |
| 499 | 00217cd41388328c93e9e9644921c4319bb03bcc | 3b337b01a1a85c2f5b49376e8c0bbdb3f521528e | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-26T13:19:23+03:00 | GitHub | noreply@github.com | 2025-09-26T10:19:23Z | | ci : create git tags for released docker images (#16008) |
| 500 | 3b337b01a1a85c2f5b49376e8c0bbdb3f521528e | a86a580a6692edd1da076d5422f944c344b4fe5e | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-26T07:53:36+02:00 | GitHub | noreply@github.com | 2025-09-26T08:53:36+03:00 | | codeowners : add danbev as owner of build-xcframework.sh [no ci] (#16268) |
| 501 | a86a580a6692edd1da076d5422f944c344b4fe5e | 0f7c69689f4e681948a6005adea9bee9c08ff903 | R0CKSTAR | yeahdongcn@gmail.com | 2025-09-26T08:56:38+08:00 | GitHub | noreply@github.com | 2025-09-26T02:56:38+02:00 | | musa: upgrade musa sdk to 4.3.0 (#16240) |
| 502 | 0f7c69689f4e681948a6005adea9bee9c08ff903 | 835b2b915c52bcabcd688d025eacff9a07b65f52 | R0CKSTAR | yeahdongcn@gmail.com | 2025-09-26T08:56:10+08:00 | GitHub | noreply@github.com | 2025-09-26T02:56:10+02:00 | | musa: fix build warnings (#15611) |
| 503 | 835b2b915c52bcabcd688d025eacff9a07b65f52 | b05a9d650f4da1e70e7e2cbe8fe44e61b39656db | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-25T19:50:28+02:00 | GitHub | noreply@github.com | 2025-09-25T19:50:28+02:00 | | model : add GroveMoE support (#15510) |
| 504 | b05a9d650f4da1e70e7e2cbe8fe44e61b39656db | 27052978e487957b3d39e3bd8fc8ee4aa304bff9 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-25T23:38:10+08:00 | GitHub | noreply@github.com | 2025-09-25T23:38:10+08:00 | | vendors: update miniaudio version (#16212) |
| 505 | 27052978e487957b3d39e3bd8fc8ee4aa304bff9 | 077c94d0caf87fbd3cf3288dbb5c0fd9670294cf | rtaluyev | taluyev@gmail.com | 2025-09-25T18:20:34+03:00 | GitHub | noreply@github.com | 2025-09-25T18:20:34+03:00 | | readme : update bindings (#16144) |
| 506 | 077c94d0caf87fbd3cf3288dbb5c0fd9670294cf | aa3ee0eb0b80efca126cedf9bcb4fb5864b46ce3 | Aman Gupta | amangupta052@gmail.com | 2025-09-25T22:35:05+08:00 | GitHub | noreply@github.com | 2025-09-25T16:35:05+02:00 | | CUDA: add a fused top-K MoE kernel (#16130) |
| 507 | aa3ee0eb0b80efca126cedf9bcb4fb5864b46ce3 | d0991da39d3c39b3980c24cdfccb9a4c95a46870 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-25T12:02:36+02:00 | GitHub | noreply@github.com | 2025-09-25T12:02:36+02:00 | | model-conversion : add embedding prompt file support (#15871) |
| 508 | d0991da39d3c39b3980c24cdfccb9a4c95a46870 | aa719c2f886c445b78113bf1e5849f1fce4124a7 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-25T11:36:47+02:00 | GitHub | noreply@github.com | 2025-09-25T11:36:47+02:00 | | server : add support for external server for tests (#16243) |
| 509 | aa719c2f886c445b78113bf1e5849f1fce4124a7 | 4cdd0bb4537f9617e9efcfef6b9454fcefe2ff08 | junchao-zhao | 68935141+junchao-loongson@users.noreply.github.com | 2025-09-25T17:22:55+08:00 | GitHub | noreply@github.com | 2025-09-25T12:22:55+03:00 | | ggml : fix loongarch lsx compilation error (#15864) |
| 510 | 4cdd0bb4537f9617e9efcfef6b9454fcefe2ff08 | b5bd037832bcb8ed3086dfe26ce9090bea989af1 | Johannes Gäßler | johannesg@5d6.de | 2025-09-25T11:12:27+02:00 | GitHub | noreply@github.com | 2025-09-25T12:12:27+03:00 | | docs: fix typo [no ci] (#16244) |
| 511 | b5bd037832bcb8ed3086dfe26ce9090bea989af1 | dfcd53f7ecb9bb897a9d752d09c59d10be47237a | Douglas Hanley | thesecretaryofwar@gmail.com | 2025-09-25T03:53:09-05:00 | GitHub | noreply@github.com | 2025-09-25T11:53:09+03:00 | | llama : add support for qwen3 reranker (#15824) |
| 512 | dfcd53f7ecb9bb897a9d752d09c59d10be47237a | 4ea00794b8c995b6deaf4bac159c1778dc27419a | Georgi Gerganov | ggerganov@gmail.com | 2025-09-25T11:30:16+03:00 | GitHub | noreply@github.com | 2025-09-25T11:30:16+03:00 | | metal : fuse NORM + MUL + ADD, support non-multiples of 4 (#16220) |
| 513 | 4ea00794b8c995b6deaf4bac159c1778dc27419a | 02a6a82ae7c7ddd1819ae26b0cc36675879aecee | Georgi Gerganov | ggerganov@gmail.com | 2025-09-25T11:29:42+03:00 | GitHub | noreply@github.com | 2025-09-25T11:29:42+03:00 | | metal : relax reorder conditions (#16216) |
| 514 | 02a6a82ae7c7ddd1819ae26b0cc36675879aecee | c498fc82fe5b83fc8c6e1627286bdc1f93caddbf | Georgi Gerganov | ggerganov@gmail.com | 2025-09-25T11:29:08+03:00 | GitHub | noreply@github.com | 2025-09-25T11:29:08+03:00 | | metal : restore im2col perf (#16219) |
| 515 | c498fc82fe5b83fc8c6e1627286bdc1f93caddbf | e7a5130a20cf45a3358308356c249734ca982143 | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-25T10:20:02+03:00 | GitHub | noreply@github.com | 2025-09-25T07:20:02Z | | rpc : use ggml logging facilities |
| 516 | e7a5130a20cf45a3358308356c249734ca982143 | bee378e0988b44bbe93b0768208080951b312363 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-25T13:06:30+08:00 | GitHub | noreply@github.com | 2025-09-25T08:06:30+03:00 | | codeowners: add ownership of zdnn backend [no ci] (#16232) |
| 517 | bee378e0988b44bbe93b0768208080951b312363 | 5fb557653b8756ac2b6c6a102b1e44abcadf552f | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-09-25T05:06:06Z | GitHub | noreply@github.com | 2025-09-25T08:06:06+03:00 | | ci: run the x64 and arm ci on the github machines instead (#16183) |
| 518 | 5fb557653b8756ac2b6c6a102b1e44abcadf552f | 4ae88d07d026e66b41e85afece74e88af54f4e66 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-25T11:36:30+08:00 | GitHub | noreply@github.com | 2025-09-25T11:36:30+08:00 | | devops: fix s390x docker release failure (#16231) |
| 519 | 4ae88d07d026e66b41e85afece74e88af54f4e66 | e789095502b337690c7616db32d7c679a5bd2533 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-25T00:25:04+08:00 | GitHub | noreply@github.com | 2025-09-25T00:25:04+08:00 | | codeowners: add ownership of zdnn backend [no ci] (#16229) |
| 520 | e789095502b337690c7616db32d7c679a5bd2533 | f2a789e33490deb483a2694b066b37e45524bb79 | Johannes Gäßler | johannesg@5d6.de | 2025-09-24T16:53:48+02:00 | GitHub | noreply@github.com | 2025-09-24T16:53:48+02:00 | | llama: print memory breakdown on exit (#15860) |
| 521 | f2a789e33490deb483a2694b066b37e45524bb79 | 3a599719673c850647e3bb911ed6d91109bb91d2 | Acly | aclysia@gmail.com | 2025-09-24T16:17:49+02:00 | GitHub | noreply@github.com | 2025-09-24T16:17:49+02:00 | | ggml : split graph allocations according to backend max buffer size (#15815) |
| 522 | 3a599719673c850647e3bb911ed6d91109bb91d2 | 63b54c81a620981be020184ab99e63a8e50e47cb | Tarek Dakhran | t.dakhran@gmail.com | 2025-09-24T13:42:26+02:00 | GitHub | noreply@github.com | 2025-09-24T13:42:26+02:00 | | model : add label for LiquidAI LFM2-2.6B model (#16204) |
| 523 | 152729f8848b692e1ae53c197ccca92ef10a523b | c0c59c1157b5aeded285a3b25d2463c866f583d3 | Uilian Ries | uilianries@gmail.com | 2025-09-24T08:53:47+02:00 | GitHub | noreply@github.com | 2025-09-24T09:53:47+03:00 | | common : add missing chrono header for common.cpp (#16211) |
| 524 | c0c59c1157b5aeded285a3b25d2463c866f583d3 | 7735706b939847df2c6c00b44c36f95a20137c61 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-24T08:53:20+02:00 | GitHub | noreply@github.com | 2025-09-24T08:53:20+02:00 | | codeowners : match all requirements files (#16214) |
| 525 | 7735706b939847df2c6c00b44c36f95a20137c61 | 4d9ea03d170cf504a40bf3a85788af84cccaee80 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-24T14:46:52+08:00 | GitHub | noreply@github.com | 2025-09-24T08:46:52+02:00 | | model-conversion : run-org-model.py fails to run on mac m1 (#16213) |
| 526 | 4d9ea03d170cf504a40bf3a85788af84cccaee80 | 8ba548dae251f8f2e3834ed5226b76b6e4f4a485 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-24T08:10:09+02:00 | GitHub | noreply@github.com | 2025-09-24T08:10:09+02:00 | | codeowners : use slash prefix for root files [no ci] (#16210) |
| 527 | 8ba548dae251f8f2e3834ed5226b76b6e4f4a485 | f505bd83ca7a43c4585ff3d59135e77eae9c793b | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-24T12:19:23+08:00 | GitHub | noreply@github.com | 2025-09-24T06:19:23+02:00 | | model-conversion : fix the make targets in the README.md (#16209) |
| 528 | f505bd83ca7a43c4585ff3d59135e77eae9c793b | 0889589dbe8c67dd518c58a57afbb02dde2dccbe | Georgi Gerganov | ggerganov@gmail.com | 2025-09-23T20:41:40+03:00 | GitHub | noreply@github.com | 2025-09-23T20:41:40+03:00 | | ci : disable AMD workflows + update NVIDIA workflows (#16200) |
| 529 | 0889589dbe8c67dd518c58a57afbb02dde2dccbe | 4e29084ba4104c4ea529fd3163bb6e76f64383df | Georgi Gerganov | ggerganov@gmail.com | 2025-09-23T13:44:25+03:00 | GitHub | noreply@github.com | 2025-09-23T13:44:25+03:00 | | ci : enable Vulkan workflow on Mac (#16194) |
| 530 | 4e29084ba4104c4ea529fd3163bb6e76f64383df | f6b4af3d04763b1e0130f5b5fce19c4bc6f83f1c | Xiangyan Sun | wishstudio@gmail.com | 2025-09-23T01:58:12-07:00 | GitHub | noreply@github.com | 2025-09-23T11:58:12+03:00 | | ggml-cpu: Respect cpumask settings (#16164) |
| 531 | f6b4af3d04763b1e0130f5b5fce19c4bc6f83f1c | 264f1b51872c125e23fa0ac1da5e2a1170de9a08 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-23T10:25:20+02:00 | GitHub | noreply@github.com | 2025-09-23T10:25:20+02:00 | | ggml : fix uninitialized is_on_grid in quantize_row_iq3_xxs_impl (#15928) |
| 532 | 264f1b51872c125e23fa0ac1da5e2a1170de9a08 | 0bc7cc715472fe28822e57f036f0746592ed2c04 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-23T14:53:05+08:00 | GitHub | noreply@github.com | 2025-09-23T14:53:05+08:00 | | zdnn: refactor codebase + add docs (#16178) |
| 533 | 0bc7cc715472fe28822e57f036f0746592ed2c04 | 4b9f4cb0f89a88de4bdf97727d0457b0c648804c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-23T08:13:22+02:00 | GitHub | noreply@github.com | 2025-09-23T09:13:22+03:00 | | codeowners : add @danbev to model-conversion example [no ci] (#16190) |
| 534 | 4b9f4cb0f89a88de4bdf97727d0457b0c648804c | 85e72271ba1ce78adf34fd8997803c991e617ca6 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-23T13:59:34+08:00 | GitHub | noreply@github.com | 2025-09-23T13:59:34+08:00 | | devops: add s390x containers (#15915) |
| 535 | 85e72271ba1ce78adf34fd8997803c991e617ca6 | 1d0125bcf1cbd7195ad0faf826a20bc7cec7d3f4 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-23T05:59:03+02:00 | GitHub | noreply@github.com | 2025-09-23T05:59:03+02:00 | | ggml-cpu : fix typo in gemm comments [no ci] (#16189) |
| 536 | 1d0125bcf1cbd7195ad0faf826a20bc7cec7d3f4 | 351f3da39c85f59d581fc184f09283da7f099a3b | Gabe Goodhart | ghart@us.ibm.com | 2025-09-22T12:40:10-06:00 | GitHub | noreply@github.com | 2025-09-22T20:40:10+02:00 | | feat: Add conversion support in GraniteHybrid for non-hybrid (all attn) (#16177) |
| 537 | 351f3da39c85f59d581fc184f09283da7f099a3b | 3ecb2f671a2f49d56357f99d135a94e841759178 | Haiyue Wang | haiyuewa@163.com | 2025-09-23T01:57:46+08:00 | GitHub | noreply@github.com | 2025-09-22T19:57:46+02:00 | | clang-tidy : disable warning about performance enum size (#16127) |
| 538 | 3ecb2f671a2f49d56357f99d135a94e841759178 | 432cf4304c5379bfbd1f3feed65ec205955b2f4b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-22T19:13:00+02:00 | GitHub | noreply@github.com | 2025-09-22T19:13:00+02:00 | | ggml : implement set_rows with i32 index (#16159) |
| 539 | 432cf4304c5379bfbd1f3feed65ec205955b2f4b | 37a23c17bdc9c99c9c6ad41168e4ced3724b72cd | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T18:20:21+03:00 | GitHub | noreply@github.com | 2025-09-22T18:20:21+03:00 | | codeowners : update + cleanup (#16174) |
| 540 | 37a23c17bdc9c99c9c6ad41168e4ced3724b72cd | 138c87ce8bd558b2cc134ada7316a3dad8eb67ac | Adrien Gallouët | angt@huggingface.co | 2025-09-22T14:13:51+02:00 | GitHub | noreply@github.com | 2025-09-22T15:13:51+03:00 | | common : enable `--offline` mode without curl support (#16137) |
| 541 | 138c87ce8bd558b2cc134ada7316a3dad8eb67ac | c6db9a10278f40b4b8d3f6931728f2cce2356daa | Quentin Bramas | quentin.bramas@gmail.com | 2025-09-22T10:53:13+02:00 | GitHub | noreply@github.com | 2025-09-22T11:53:13+03:00 | | webui : fix handling incomplete chunks (#16107) |
| 542 | c6db9a10278f40b4b8d3f6931728f2cce2356daa | d05affbab7d1364fa7ac06c03cd33f03235ae840 | GideonSerf | gdserf.gs@gmail.com | 2025-09-22T10:49:58+02:00 | GitHub | noreply@github.com | 2025-09-22T11:49:58+03:00 | | embedding : fix typos in README (#16171) |
| 543 | d05affbab7d1364fa7ac06c03cd33f03235ae840 | 4f324a556caa5be0be2261604ef4aab0faf36eb8 | Haiyue Wang | haiyuewa@163.com | 2025-09-22T16:48:42+08:00 | GitHub | noreply@github.com | 2025-09-22T11:48:42+03:00 | | common : remove unused local variables (#16140) |
| 544 | 4f324a556caa5be0be2261604ef4aab0faf36eb8 | a71ae3ba7a857394103f12e30b387c48242c84f2 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T11:12:37+03:00 | GitHub | noreply@github.com | 2025-09-22T11:12:37+03:00 | | ggml : extend ggml_can_fuse to work with non-sequential nodes (#16123) |
| 545 | a71ae3ba7a857394103f12e30b387c48242c84f2 | 05a2458121f9cdcc98ea6ffeeab6f23023136266 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T11:12:09+03:00 | GitHub | noreply@github.com | 2025-09-22T11:12:09+03:00 | | ggml : add ggml_op_is_empty (#16122) |
| 546 | 05a2458121f9cdcc98ea6ffeeab6f23023136266 | 96fdca043be8fa733bbd59868a75517f92633376 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-22T15:10:58+07:00 | GitHub | noreply@github.com | 2025-09-22T11:10:58+03:00 | | codeowners : update ownership for @ngxson and @allozuar (#16128) |
| 547 | 96fdca043be8fa733bbd59868a75517f92633376 | b2d980fce063fa3591c59ad50b946c8da66c294f | Shin-myoung-serp | relent95@naver.com | 2025-09-22T17:04:01+09:00 | GitHub | noreply@github.com | 2025-09-22T10:04:01+02:00 | | Vulkan: add conv_transpose_2d operation (#16022) |
| 548 | b2d980fce063fa3591c59ad50b946c8da66c294f | 5c6106a696af2b00ce01d14ba525979d5a32b199 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-22T09:59:05+02:00 | GitHub | noreply@github.com | 2025-09-22T10:59:05+03:00 | | codeowners : claim responsibility for ci, models, gguf-py and convert (#16124) |
| 549 | 5c6106a696af2b00ce01d14ba525979d5a32b199 | ec65fb52f0cc890e5085a6c5995ab08a156265fb | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T10:58:02+03:00 | GitHub | noreply@github.com | 2025-09-22T10:58:02+03:00 | | contrib : update roles (#16113) |
| 550 | ec65fb52f0cc890e5085a6c5995ab08a156265fb | 1d660d2fae42ea2e1d3569638e722bf7a37b6b19 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T10:16:05+03:00 | GitHub | noreply@github.com | 2025-09-22T10:16:05+03:00 | | ci : remove vulkaninfo calls (#16169) |
| 551 | 1d660d2fae42ea2e1d3569638e722bf7a37b6b19 | a20d810d79adfbf82e0eb703af6f27a0e2d1a539 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T09:11:39+03:00 | GitHub | noreply@github.com | 2025-09-22T09:11:39+03:00 | | ci : use smaller model (#16168) |
| 552 | a20d810d79adfbf82e0eb703af6f27a0e2d1a539 | 4d0a7cbc617e384fc355077a304c883b5c7d4fb6 | Jeff Bolz | jbolz@nvidia.com | 2025-09-22T00:37:17-05:00 | GitHub | noreply@github.com | 2025-09-22T07:37:17+02:00 | | vulkan: add RTE variants of exp shader (#16165) |
| 553 | 4d0a7cbc617e384fc355077a304c883b5c7d4fb6 | 9073a73d82a916cea0809de225ef5175c3a86e91 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-22T08:31:40+03:00 | GitHub | noreply@github.com | 2025-09-22T08:31:40+03:00 | | ci : adjust params for less runtime (#16167) |
| 554 | 9073a73d82a916cea0809de225ef5175c3a86e91 | 51f5a45fbe575dcd54bdd2a339ef8e8424d1c12a | Ruben Ortlam | picard12@live.de | 2025-09-22T07:22:43+02:00 | GitHub | noreply@github.com | 2025-09-22T07:22:43+02:00 | | vulkan: vec dot matrix multiplication fix (#16151) |
| 555 | 51f5a45fbe575dcd54bdd2a339ef8e8424d1c12a | c4510dc9374e17dcb8726902ab5216067a92b3d3 | lhez | lih@qti.qualcomm.com | 2025-09-21T16:42:10-07:00 | GitHub | noreply@github.com | 2025-09-21T16:42:10-07:00 | | opencl: fix concat crash on win arm64 with Adreno (#15944) |
| 556 | c4510dc9374e17dcb8726902ab5216067a92b3d3 | da30ab5f8696cabb2d4620cdc0aa41a298c54fd6 | lhez | lih@qti.qualcomm.com | 2025-09-21T14:48:44-07:00 | GitHub | noreply@github.com | 2025-09-21T14:48:44-07:00 | | opencl: initial `q8_0` mv support (#15732) |
| 557 | da30ab5f8696cabb2d4620cdc0aa41a298c54fd6 | 28baac9c9f491c872e2c37762d3bd90446b005e9 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-21T19:00:27+03:00 | GitHub | noreply@github.com | 2025-09-21T19:00:27+03:00 | | ci : add label for the RISC-V runner (#16150) |
| 558 | 28baac9c9f491c872e2c37762d3bd90446b005e9 | 1eeb523c3e0c7ffbd59469f5463dcbdecba3535e | Georgi Gerganov | ggerganov@gmail.com | 2025-09-21T16:50:45+03:00 | GitHub | noreply@github.com | 2025-09-21T16:50:45+03:00 | | ci : migrate ggml ci to self-hosted runners (#16116) |
| 559 | 1eeb523c3e0c7ffbd59469f5463dcbdecba3535e | 5bb4a3edec297e74b0f7bd4ed5d0fdd12e28d858 | Giuseppe Scrivano | gscrivan@redhat.com | 2025-09-21T08:31:55+02:00 | GitHub | noreply@github.com | 2025-09-21T08:31:55+02:00 | | vulkan: optimize UMA buffer operations and fix driver hangs (#16059) |
| 560 | 5bb4a3edec297e74b0f7bd4ed5d0fdd12e28d858 | 7f766929ca8e8e01dcceb1c526ee584f7e5e1408 | Jeff Bolz | jbolz@nvidia.com | 2025-09-21T01:23:37-05:00 | GitHub | noreply@github.com | 2025-09-21T08:23:37+02:00 | | vulkan: fix validation error about VK_PIPELINE_CREATE_CAPTURE_STATISTICS_BIT_KHR (#16086) |
| 561 | 7f766929ca8e8e01dcceb1c526ee584f7e5e1408 | 405921dcefdd4e90ed948a4bf179007c2fa92b2d | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T12:55:47+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T13:02:14+03:00 | | sync : ggml |
| 562 | 405921dcefdd4e90ed948a4bf179007c2fa92b2d | fa6383ca7e7ccb8ca3bdfeb37e348ddc4113aa26 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-16T06:16:52+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T13:02:14+03:00 | | ggml : introduce semantic versioning (ggml/1336) |
| 563 | fa6383ca7e7ccb8ca3bdfeb37e348ddc4113aa26 | 803dac2e48ef3ba26a504eb27c4e77ec2d21f7d0 | Gregor Jasny | gjasny@googlemail.com | 2025-09-10T17:21:11+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-20T13:02:14+03:00 | | CUDA : conditionally add cuda architectures (ggml/1341) |
| 564 | 803dac2e48ef3ba26a504eb27c4e77ec2d21f7d0 | 459c0c2c1a400f960d7b8e8d94d31a8426f80986 | Ruben Ortlam | picard12@live.de | 2025-09-20T10:42:56+02:00 | GitHub | noreply@github.com | 2025-09-20T10:42:56+02:00 | | vulkan: use vec dot for matrix matrix multiplications (#16056) |
| 565 | 459c0c2c1a400f960d7b8e8d94d31a8426f80986 | be79d9fdd95ab8955527c4aaa67b90e8b9516718 | Benni | 73313922+BenjaminBruenau@users.noreply.github.com | 2025-09-20T07:56:30+02:00 | GitHub | noreply@github.com | 2025-09-20T07:56:30+02:00 | | server: fix SSE and OpenAI compatibility for error messages when streaming (#16109) |
| 566 | be79d9fdd95ab8955527c4aaa67b90e8b9516718 | f432d8d83e7407073634c5e4fd81a3d23a10827f | ssweens | 1149151+ssweens@users.noreply.github.com | 2025-09-19T15:15:21-07:00 | GitHub | noreply@github.com | 2025-09-20T00:15:21+02:00 | | llama-bench: add --devices and --list-devices support (#16039) |
| 567 | f432d8d83e7407073634c5e4fd81a3d23a10827f | 4067f07fc5aa5722b87db303b561597004696f6c | shun095 | 8069181+shun095@users.noreply.github.com | 2025-09-20T00:57:30+09:00 | GitHub | noreply@github.com | 2025-09-19T09:57:30-06:00 | | chat: Fix streaming parser for granite models (#15682) |
| 568 | 4067f07fc5aa5722b87db303b561597004696f6c | 4b8560ab56fdd9819358b47c338bbc8ec357c57e | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-19T09:52:27+02:00 | GitHub | noreply@github.com | 2025-09-19T09:52:27+02:00 | | feat: Improve mobile UI for Settings Dialog (#16084) |
| 569 | 4b8560ab56fdd9819358b47c338bbc8ec357c57e | 0dd58b6877f3dc106593d6bc68d98305f553c566 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-19T13:02:51+07:00 | GitHub | noreply@github.com | 2025-09-19T13:02:51+07:00 | | chat : fix build on arm64 (#16101) |
| 570 | 0dd58b6877f3dc106593d6bc68d98305f553c566 | 69ffd891631befa9e6b485fd646a16dab4f2c007 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-19T11:31:56+07:00 | GitHub | noreply@github.com | 2025-09-19T06:31:56+02:00 | | ggml : refactor forward_dup for cpu backend (#16062) |
| 571 | 69ffd891631befa9e6b485fd646a16dab4f2c007 | 246c0d9c795fb25da8268625d597580155a3672f | Adrien Gallouët | angt@huggingface.co | 2025-09-18T23:07:26+02:00 | GitHub | noreply@github.com | 2025-09-18T23:07:26+02:00 | | ggml-amx : fix ggml_amx_init() on generic Linux (#16049) |
| 572 | 246c0d9c795fb25da8268625d597580155a3672f | 3edd87cd055a45d885fa914d879d36d33ecfc3e1 | Adrien Gallouët | angt@huggingface.co | 2025-09-18T23:07:18+02:00 | GitHub | noreply@github.com | 2025-09-18T23:07:18+02:00 | | cmake : fix static linking for OpenMP on Unix-like systems (#16031) |
| 573 | 3edd87cd055a45d885fa914d879d36d33ecfc3e1 | c0b45097c33e2667a94444f08cc9e36bec0a5e2e | Shawn Gu | shawngu@qti.qualcomm.com | 2025-09-18T12:03:34-07:00 | GitHub | noreply@github.com | 2025-09-18T12:03:34-07:00 | | opencl: optimize mxfp4 kernels (#16037) |
| 574 | c0b45097c33e2667a94444f08cc9e36bec0a5e2e | 38dbdf4c057515ccea9bec0ca2518f86d5e4d28e | Jeff Bolz | jbolz@nvidia.com | 2025-09-18T13:46:17-05:00 | GitHub | noreply@github.com | 2025-09-18T13:46:17-05:00 | | rename optimize_graph to graph_optimize (#16082) |
| 575 | 38dbdf4c057515ccea9bec0ca2518f86d5e4d28e | 368560a1e3b9a3bc83af741b0b2bc9e46fb420d2 | Bowen Han | fancycode@gmail.com | 2025-09-18T11:26:03-07:00 | GitHub | noreply@github.com | 2025-09-18T20:26:03+02:00 | | CUDA: Optimize PAD_REFLECT_1D (#15957) |
| 576 | 368560a1e3b9a3bc83af741b0b2bc9e46fb420d2 | 4ca088b036313c2d8e682f4cfeb7c29edd85d0b9 | Johannes Gäßler | johannesg@5d6.de | 2025-09-18T19:28:32+02:00 | GitHub | noreply@github.com | 2025-09-18T19:28:32+02:00 | | CUDA: fix compilation on CC 6.0 (#16091) |
| 577 | 4ca088b036313c2d8e682f4cfeb7c29edd85d0b9 | 703f9e32c4eb3166f8d63007c26e31a1466c21af | Eric Curtin | eric.curtin@docker.com | 2025-09-18T16:22:50+01:00 | GitHub | noreply@github.com | 2025-09-18T16:22:50+01:00 | | Add resumable downloads for llama-server model loading (#15963) |
| 578 | 703f9e32c4eb3166f8d63007c26e31a1466c21af | ad6bd9083bc346afd4ecd2fbb83e0306c0141c99 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-18T16:28:41+03:00 | GitHub | noreply@github.com | 2025-09-18T16:28:41+03:00 | | metal : use function constants for mul_mv_ext kernels (#16074) |
| 579 | ad6bd9083bc346afd4ecd2fbb83e0306c0141c99 | 2b6b55a59f590a1e1eb2dcd09d5b8b6feb9cc748 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-18T13:28:22+02:00 | GitHub | noreply@github.com | 2025-09-18T13:28:22+02:00 | | cuda : add missing F32<->I32 entries in ggml_cuda_cpy_fn (#16060) |
| 580 | 2b6b55a59f590a1e1eb2dcd09d5b8b6feb9cc748 | e58174cecbc45bf79bf653cd2c984395940c6ef4 | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-18T13:36:57+03:00 | GitHub | noreply@github.com | 2025-09-18T10:36:57Z | | server : include usage statistics only when user request them (#16052) |
| 581 | e58174cecbc45bf79bf653cd2c984395940c6ef4 | b213fce89bee8cb56b587b91e15a4278f8ed0180 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-18T12:47:56+03:00 | GitHub | noreply@github.com | 2025-09-18T12:47:56+03:00 | | llama : bump max seq limit from 64 to 256 (#15916) |
| 582 | b213fce89bee8cb56b587b91e15a4278f8ed0180 | e00f3fd8fff2cf5a8c8c9f475034bd089c8bcce4 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-18T12:33:45+03:00 | GitHub | noreply@github.com | 2025-09-18T12:33:45+03:00 | | metal : improve F32, F16 and BF16 mat-vec multiplication (#16057) |
| 583 | e00f3fd8fff2cf5a8c8c9f475034bd089c8bcce4 | f2f28380ea68939c0481df3e4216b698824a59df | Jhen-Jie Hong | iainst0409@gmail.com | 2025-09-18T15:06:48+08:00 | GitHub | noreply@github.com | 2025-09-18T10:06:48+03:00 | | metal : avoid call free for non-owned buffer (#16067) |
| 584 | f2f28380ea68939c0481df3e4216b698824a59df | 62c3b645c56ad5288aa401901f043b050edac5a2 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-18T10:03:24+03:00 | GitHub | noreply@github.com | 2025-09-18T10:03:24+03:00 | | metal : handle nil cv during pipeline creation (#16065) |
| 585 | 62c3b645c56ad5288aa401901f043b050edac5a2 | d304f459d8619399eb232b1e3689379306f439d3 | Chenguang Li | 757486878@qq.com | 2025-09-18T09:26:33+08:00 | GitHub | noreply@github.com | 2025-09-18T09:26:33+08:00 | | CANN: Remove print (#16044) |
| 586 | d304f459d8619399eb232b1e3689379306f439d3 | 0320ac5264279d74f8ee91bafa6c90e9ab9bbb91 | Reese Levine | reeselevine1@gmail.com | 2025-09-17T13:09:40-07:00 | GitHub | noreply@github.com | 2025-09-17T13:09:40-07:00 | | GGML WebGPU: Support for ADD, MUL, RMS_NORM, GET_ROWS operators (#16018) |
| 587 | 0320ac5264279d74f8ee91bafa6c90e9ab9bbb91 | a7a98e0fffed794396b3fbad4dcdbbc184963645 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-17T20:38:12+03:00 | GitHub | noreply@github.com | 2025-09-17T20:38:12+03:00 | | metal : refactor + optimize v2 (#15995) |
| 588 | a7a98e0fffed794396b3fbad4dcdbbc184963645 | 8f8f2274ee3601fecf6e2d57b52f701c81bede21 | Aleksander Grygier | aleksander.grygier@gmail.com | 2025-09-17T19:29:13+02:00 | GitHub | noreply@github.com | 2025-09-17T19:29:13+02:00 | | SvelteKit-based WebUI (#14839) |
| 589 | 8f8f2274ee3601fecf6e2d57b52f701c81bede21 | c959b676be29e93f8dbc3bd6056ceba812a9eb72 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-18T00:18:21+07:00 | GitHub | noreply@github.com | 2025-09-17T19:18:21+02:00 | | convert : add Llama4ForCausalLM (#16042) |
| 590 | c959b676be29e93f8dbc3bd6056ceba812a9eb72 | cd08fc3ecc0264b4414b68af3874a6c689ed60c1 | Johannes Gäßler | johannesg@5d6.de | 2025-09-17T15:32:42+02:00 | GitHub | noreply@github.com | 2025-09-17T15:32:42+02:00 | | CUDA: fix FA occupancy, optimize tile kernel (#15982) |
| 591 | cd08fc3ecc0264b4414b68af3874a6c689ed60c1 | cb5bb6cc05119c24e7711ca2956cd0e6d409d396 | David Ribeiro Alves | davidralves@gmail.com | 2025-09-17T01:08:02-07:00 | GitHub | noreply@github.com | 2025-09-17T11:08:02+03:00 | | common : Fix corrupted memory error on json grammar initialization (#16038) |
| 592 | cb5bb6cc05119c24e7711ca2956cd0e6d409d396 | a91d035b901e8a9edf810f63d130ee49adf27be2 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-09-17T07:35:37Z | GitHub | noreply@github.com | 2025-09-17T09:35:37+02:00 | | vulkan: automatically remove unsupported devices (#15976) |
| 593 | a91d035b901e8a9edf810f63d130ee49adf27be2 | 745cbcf2fe1eb88f8db615ac622f0b944d924ad6 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-17T09:34:09+02:00 | GitHub | noreply@github.com | 2025-09-17T09:34:09+02:00 | | ci : revert back to macos-13 for macOS-latest-cmake-x64 (#16040) |
| 594 | 745cbcf2fe1eb88f8db615ac622f0b944d924ad6 | 1cbd80f8cf80a817715b1ccc5680fe2a3c5172c8 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-17T15:30:55+08:00 | GitHub | noreply@github.com | 2025-09-17T09:30:55+02:00 | | llama-quant : fix the verification of attention layers for encoder-decoder models (#16023) |
| 595 | 1cbd80f8cf80a817715b1ccc5680fe2a3c5172c8 | 85286f354813056f6c835046c0acfa3bf6ba9432 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-17T15:29:00+08:00 | GitHub | noreply@github.com | 2025-09-17T10:29:00+03:00 | | examples : support encoder-decoder models in the simple example (#16002) |
| 596 | 85286f354813056f6c835046c0acfa3bf6ba9432 | d5fabe3682de515fd09d6c981f7a0d1b75614455 | Shane A | shanea@allenai.org | 2025-09-17T00:01:58-07:00 | GitHub | noreply@github.com | 2025-09-17T09:01:58+02:00 | | model : add OLMo3 support (#16015) |
| 597 | d5fabe3682de515fd09d6c981f7a0d1b75614455 | 8ff206097c2bf3ca1c7aa95f9d6db779fc7bdd68 | Chenguang Li | 757486878@qq.com | 2025-09-17T14:33:08+08:00 | GitHub | noreply@github.com | 2025-09-17T14:33:08+08:00 | | CANN: Optimize ggml_cann_set_device (#15935) |
| 598 | 8ff206097c2bf3ca1c7aa95f9d6db779fc7bdd68 | 77475530b8bbea3bf578632507e1284cdfe2c8c0 | jacekpoplawski | 67507230+jacekpoplawski@users.noreply.github.com | 2025-09-16T16:17:08+02:00 | GitHub | noreply@github.com | 2025-09-16T16:17:08+02:00 | | llama-bench: add --n-cpu-moe support (#15952) |
| 599 | 77475530b8bbea3bf578632507e1284cdfe2c8c0 | 3913f8730ec6d6245480affc30ae3049107956f4 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-16T15:27:52+02:00 | GitHub | noreply@github.com | 2025-09-16T15:27:52+02:00 | | ci : use macos-latest for arm64 webgpu build (#16029) |
| 600 | 3913f8730ec6d6245480affc30ae3049107956f4 | 76888d202ed2b835ae19ea9f9db6baf39e419297 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-16T15:25:57+02:00 | GitHub | noreply@github.com | 2025-09-16T15:25:57+02:00 | | ggml : fix padding in timestep embedding kernels (#15932) |
| 601 | 76888d202ed2b835ae19ea9f9db6baf39e419297 | f1fbffb5c0b34b2a68febb7da3fd0f8333f1ed4c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-16T13:41:38+02:00 | GitHub | noreply@github.com | 2025-09-16T13:41:38+02:00 | | ci : upload xcframework artifact from ios-xcode-build job (#16010) |
| 602 | f1fbffb5c0b34b2a68febb7da3fd0f8333f1ed4c | 51abc96bdc52ba8cd6ad78dcf12ed9a041d7b442 | Bowen Han | fancycode@gmail.com | 2025-09-15T23:59:19-07:00 | GitHub | noreply@github.com | 2025-09-16T08:59:19+02:00 | | fix: apply clang-format to CUDA macros (#16017) |
| 603 | 51abc96bdc52ba8cd6ad78dcf12ed9a041d7b442 | 07808ebb07e2b1aa19032705e332679ddf967614 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-16T05:57:16+02:00 | GitHub | noreply@github.com | 2025-09-16T05:57:16+02:00 | | ci : update macos-latest* jobs to use macos-latest (#15938) |
| 604 | 07808ebb07e2b1aa19032705e332679ddf967614 | 6d758839ff741d4966ca92b7f801b7a8b5b96364 | Yuri Khrustalev | ykhrustalev@users.noreply.github.com | 2025-09-15T22:54:44-04:00 | GitHub | noreply@github.com | 2025-09-16T09:54:44+07:00 | | cmake : Do not install tools on iOS targets (#15903) |
| 605 | 6d758839ff741d4966ca92b7f801b7a8b5b96364 | 3d4053f77f0f78ee2b791088c02af653ebee42dd | Aman Gupta | amangupta052@gmail.com | 2025-09-16T10:38:28+08:00 | GitHub | noreply@github.com | 2025-09-16T10:38:28+08:00 | | Add LLaDA-7b-MoE diffusion model (#16003) |
| 606 | 3d4053f77f0f78ee2b791088c02af653ebee42dd | dc381aa9a6dc45f00673471d34b8bddd30e77570 | Jake Karnes | jake.karnes@gmail.com | 2025-09-15T16:28:31-06:00 | GitHub | noreply@github.com | 2025-09-16T00:28:31+02:00 | | CUDA: fix im2col_3d to respect non-contiguous inputs (views) (#15956) |
| 607 | dc381aa9a6dc45f00673471d34b8bddd30e77570 | 10d197409bd9537ff302ad09966fe406882fef9d | Diego Devesa | slarengh@gmail.com | 2025-09-15T14:38:52-07:00 | GitHub | noreply@github.com | 2025-09-15T23:38:52+02:00 | | docker : enable rocWMMA in ROCm images, add gfx1151 (#15997) |
| 608 | 10d197409bd9537ff302ad09966fe406882fef9d | b907255f4bd169b0dc7dca9553b4c54af5170865 | Diego Devesa | slarengh@gmail.com | 2025-09-15T14:38:42-07:00 | GitHub | noreply@github.com | 2025-09-15T23:38:42+02:00 | | releases : switch to rocWMMA develop branch, add gfx1151 (#15992) |
| 609 | b907255f4bd169b0dc7dca9553b4c54af5170865 | 28c39da7c645185ade5436767929d7ec33006033 | yael-works | 106673277+yael-works@users.noreply.github.com | 2025-09-15T19:51:35+03:00 | GitHub | noreply@github.com | 2025-09-15T18:51:35+02:00 | | SYCL: Add COUNT_EQUAL operator support (#15991) |
| 610 | 28c39da7c645185ade5436767929d7ec33006033 | 106220562aca42b6738b8f51acfce0db1b8a2fb6 | Nikolay Popov | 131475237+npopov-vst@users.noreply.github.com | 2025-09-15T13:08:30+03:00 | GitHub | noreply@github.com | 2025-09-15T11:08:30+01:00 | | llama-run: Fix model download on Windows (#15988) |
| 611 | 106220562aca42b6738b8f51acfce0db1b8a2fb6 | a68f31edd71cc39141113f05f7133a3e9ece8c61 | Aman Gupta | amangupta052@gmail.com | 2025-09-15T17:35:11+08:00 | GitHub | noreply@github.com | 2025-09-15T17:35:11+08:00 | | CUDA: some micro-optimizations in mmf.cuh for mul_mat_id (#15926) |
| 612 | a68f31edd71cc39141113f05f7133a3e9ece8c61 | b8e09f08b9a91c0401bc67d17a17c90756420346 | ddh0 | chemist-mulches-39@icloud.com | 2025-09-15T02:54:57-05:00 | GitHub | noreply@github.com | 2025-09-15T09:54:57+02:00 | | fix KLD percentile output (#15999) |
| 613 | b8e09f08b9a91c0401bc67d17a17c90756420346 | 6c019cb04e86e2dacfe62ce7666c64e9717dde1f | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-14T23:00:59+02:00 | GitHub | noreply@github.com | 2025-09-14T23:00:59+02:00 | | model : add grok-2 support (#15539) |
| 614 | 6c019cb04e86e2dacfe62ce7666c64e9717dde1f | 9dcd200d57bc6f05a59a6a8df361d5d183af4124 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-14T21:17:04+02:00 | GitHub | noreply@github.com | 2025-09-14T21:17:04+02:00 | | server : only attempt to enable thinking if using jinja (#15967) |
| 615 | 9dcd200d57bc6f05a59a6a8df361d5d183af4124 | 0fa154e3502e940df914f03b41475a2b80b985b0 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-14T22:02:32+03:00 | GitHub | noreply@github.com | 2025-09-14T22:02:32+03:00 | | metal : remove memory pools (#15966) |
| 616 | 0fa154e3502e940df914f03b41475a2b80b985b0 | 261e6a20ffdb79c4875e674b4f6b514bc73cff8f | Adam | channeladam@users.noreply.github.com | 2025-09-15T04:43:54+10:00 | GitHub | noreply@github.com | 2025-09-14T20:43:54+02:00 | | rocm.Dockerfile: added gfx1200,gfx1201 architectures to support AMD Radeon RX 9000 series (#15994) |
| 617 | 261e6a20ffdb79c4875e674b4f6b514bc73cff8f | a0e13dcbe5bae7025660349ef3e4ead060e507f2 | Ruben Ortlam | picard12@live.de | 2025-09-14T16:56:28+02:00 | GitHub | noreply@github.com | 2025-09-14T16:56:28+02:00 | | Vulkan: Clean up mul_mm shader (#15987) |
| 618 | a0e13dcbe5bae7025660349ef3e4ead060e507f2 | a14bd350141fb42b8bf2dd2342cebc27bfdce399 | lcy | lcy0321@users.noreply.github.com | 2025-09-14T22:20:35+08:00 | GitHub | noreply@github.com | 2025-09-14T07:20:35-07:00 | | build: fix the build failures of Windows HIP release job (#15984) |
| 619 | a14bd350141fb42b8bf2dd2342cebc27bfdce399 | 918b26f197f55d5d562446dfc876d0e637929d07 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-14T15:33:22+03:00 | GitHub | noreply@github.com | 2025-09-14T15:33:22+03:00 | | metal : fix kernel requirements (#15983) |
| 620 | 918b26f197f55d5d562446dfc876d0e637929d07 | 9ecb88434644c865232bb665d2f6f05049fc6456 | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-14T12:28:18+03:00 | GitHub | noreply@github.com | 2025-09-14T12:28:18+03:00 | | rpc : fix regression when --device is used (#15981) |
| 621 | 9ecb88434644c865232bb665d2f6f05049fc6456 | d1c6f11f47bae063a68a6be8e4830c060a11b7bd | Diego Devesa | slarengh@gmail.com | 2025-09-14T02:21:59-07:00 | GitHub | noreply@github.com | 2025-09-14T02:21:59-07:00 | | releases : update ROCM, add gfx1200, gfx1201, gfx1151 (#15972) |
| 622 | d1c6f11f47bae063a68a6be8e4830c060a11b7bd | 6380d6a3e709cb02a8695afdd96b40e674477332 | Radoslav Gerganov | rgerganov@gmail.com | 2025-09-14T12:10:07+03:00 | GitHub | noreply@github.com | 2025-09-14T12:10:07+03:00 | | doc : update documentation for --tensor-split (#15980) |
| 623 | 6380d6a3e709cb02a8695afdd96b40e674477332 | aa0c461efe3603639af1a1defed2438d9c16ca0f | Aaron Teo | aaron.teo1@ibm.com | 2025-09-14T13:37:03+08:00 | GitHub | noreply@github.com | 2025-09-14T13:37:03+08:00 | | ggml-zdnn: rm user mapped buffers (#15965) |
| 624 | aa0c461efe3603639af1a1defed2438d9c16ca0f | b9c9c9f789cd57fb0b28b25223305613cd90fa10 | Jeff Bolz | jbolz@nvidia.com | 2025-09-13T16:29:43+01:00 | GitHub | noreply@github.com | 2025-09-13T17:29:43+02:00 | | vulkan: fix failing dequant shaders (#15862) |
| 625 | b9c9c9f789cd57fb0b28b25223305613cd90fa10 | 50f4281a6f5c3a5d68bdeb12f904fa01e0e2ba91 | Jeff Bolz | jbolz@nvidia.com | 2025-09-13T16:23:30+01:00 | GitHub | noreply@github.com | 2025-09-13T17:23:30+02:00 | | vulkan: initialize vulkan-hpp to allow using extension function pointers (#15705) |
| 626 | 50f4281a6f5c3a5d68bdeb12f904fa01e0e2ba91 | 55758b00cae1451a0d789c50a5eb8c9bee42b9a0 | Diego Devesa | slarengh@gmail.com | 2025-09-13T07:49:49-07:00 | GitHub | noreply@github.com | 2025-09-13T16:49:49+02:00 | | llama : allow using iGPUs with --device (#15951) |
| 627 | 55758b00cae1451a0d789c50a5eb8c9bee42b9a0 | f161463a54d9f93d41246286aa4a9569a91d804d | Georgi Gerganov | ggerganov@gmail.com | 2025-09-13T16:24:22+03:00 | GitHub | noreply@github.com | 2025-09-13T16:24:22+03:00 | | metal : refactor kernel loading (#15964) |
| 628 | f161463a54d9f93d41246286aa4a9569a91d804d | 84d7b2fca11d1be118ce776f6d72a486c4883b74 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-13T13:54:28+03:00 | GitHub | noreply@github.com | 2025-09-13T13:54:28+03:00 | | metal : allow ops to run concurrently (#15929) |
| 629 | 84d7b2fca11d1be118ce776f6d72a486c4883b74 | 40be51152d4dc2d47444a4ed378285139859895b | Georgi Gerganov | ggerganov@gmail.com | 2025-09-13T12:45:04+03:00 | GitHub | noreply@github.com | 2025-09-13T12:45:04+03:00 | | metal : fix memory leaks (#15962) |
| 630 | 40be51152d4dc2d47444a4ed378285139859895b | 4bf5549269d99bac936a65892da2312cc71b2421 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-13T02:39:52+08:00 | GitHub | noreply@github.com | 2025-09-13T02:39:52+08:00 | | ggml-zdnn: fix #15414, activate FP16 and BF16 acceleration and incorrect zTensor free (#15839) |
| 631 | 4bf5549269d99bac936a65892da2312cc71b2421 | f4e664f838cc65d89fa845c48c372e43852112e4 | Eric Curtin | eric.curtin@docker.com | 2025-09-12T16:31:50+01:00 | GitHub | noreply@github.com | 2025-09-12T16:31:50+01:00 | | Add docker protocol support for llama-server model loading (#15790) |
| 632 | f4e664f838cc65d89fa845c48c372e43852112e4 | f088b6a84f6aed0abb619dbb9d375e726fc888a6 | Haiyue Wang | haiyuewa@163.com | 2025-09-12T23:16:32+08:00 | GitHub | noreply@github.com | 2025-09-12T18:16:32+03:00 | | context : remove redundant explicit casting to the same type (#15948) |
| 633 | f088b6a84f6aed0abb619dbb9d375e726fc888a6 | 304ac5693d1e2124f83e0584bc5eea6311d3d3b4 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-12T17:02:55+03:00 | GitHub | noreply@github.com | 2025-09-12T17:02:55+03:00 | | server : adjust prompt similarity thold + add logs (#15913) |
| 634 | 304ac5693d1e2124f83e0584bc5eea6311d3d3b4 | 6c88ad8fa741a2182a2dcf8c5353ae4b5c66e656 | Ruben Ortlam | picard12@live.de | 2025-09-12T13:24:21+02:00 | GitHub | noreply@github.com | 2025-09-12T13:24:21+02:00 | | Vulkan iGPU device selection overhaul and PCI ID API support (#15947) |
| 635 | 6c88ad8fa741a2182a2dcf8c5353ae4b5c66e656 | 704d90c987cdf00751567b2088c4e54742aa2d3f | Mathieu Baudier | mbaudier@argeo.org | 2025-09-12T09:06:20+02:00 | GitHub | noreply@github.com | 2025-09-12T09:06:20+02:00 | | vulkan: Make device memory check more portable (#15939) |
| 636 | 360d6533db39e11577afe9b0aece20c6b5ddaf1f | 0e6ff0046f4a2983b2c77950aa75960fe4b4f0e2 | Diego Devesa | slarengh@gmail.com | 2025-09-11T13:47:38-07:00 | GitHub | noreply@github.com | 2025-09-11T22:47:38+02:00 | | ggml-backend : add GGML_BACKEND_DEVICE_TYPE_IGPU device type (#15797) |
| 637 | 0e6ff0046f4a2983b2c77950aa75960fe4b4f0e2 | df082f56309073ecf885eceaa21b86e8a487e61b | Johannes Gäßler | johannesg@5d6.de | 2025-09-11T21:19:58+02:00 | GitHub | noreply@github.com | 2025-09-11T21:19:58+02:00 | | CUDA: larger SRAM reads for tile FA, AMD FP16 dot (#15927) |
| 638 | df082f56309073ecf885eceaa21b86e8a487e61b | 24a6734daf6932ff29ba8c1ff0245c51d76f783e | ddh0 | chemist-mulches-39@icloud.com | 2025-09-11T12:12:34-05:00 | GitHub | noreply@github.com | 2025-09-11T19:12:34+02:00 | | nitpick : correct MB to MiB (#15934) |
| 639 | 24a6734daf6932ff29ba8c1ff0245c51d76f783e | 2b3efea9a4d91216850856fbb77075db26f6a6eb | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-11T15:39:12+02:00 | GitHub | noreply@github.com | 2025-09-11T14:39:12+01:00 | | ggml-cpu : add check for ARM MATMUL_INT8/i8mm support (#15922) |
| 640 | 2b3efea9a4d91216850856fbb77075db26f6a6eb | c0389dba43d50695f9d3f57dd1f1a14cbefc100c | Charles Xu | charles.xu@arm.com | 2025-09-11T12:45:40+02:00 | GitHub | noreply@github.com | 2025-09-11T12:45:40+02:00 | | kleidiai: fix GGML_ASSERT(*cur_backend_id != -1) failed (#15614) |
| 641 | c0389dba43d50695f9d3f57dd1f1a14cbefc100c | 00681dfc16ba4cebb9c7fbd2cf2656e06a0692a4 | hipudding | huafengchun@gmail.com | 2025-09-11T15:59:37+08:00 | GitHub | noreply@github.com | 2025-09-11T15:59:37+08:00 | | CANN: Disable acl_graph for prefill stage (#15933) |
| 642 | 00681dfc16ba4cebb9c7fbd2cf2656e06a0692a4 | 4f658855fa8f2e42b7ed9a5b298fa39a2e39b096 | Oliver Simons | osimons@nvidia.com | 2025-09-10T22:04:03+02:00 | GitHub | noreply@github.com | 2025-09-10T22:04:03+02:00 | | CUDA: Add `fastdiv` to `k_bin_bcast*`, giving 1-3% E2E performance (#15872) |
| 643 | 4f658855fa8f2e42b7ed9a5b298fa39a2e39b096 | 6ab397e12ba8e9f776341cdae68f7ffb2f8d2cde | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-11T02:51:51+08:00 | GitHub | noreply@github.com | 2025-09-10T20:51:51+02:00 | | llama : support T5 models with unequal number of encoder-decoder layers (#15909) |
| 644 | 6ab397e12ba8e9f776341cdae68f7ffb2f8d2cde | 9de447d94e1ae9d1a36e5a2e5bf47483352c0d9c | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-10T19:08:59+02:00 | GitHub | noreply@github.com | 2025-09-10T19:08:59+02:00 | | graph : support non-contiguous Q in build_attn_mha (#15908) |
| 645 | 9de447d94e1ae9d1a36e5a2e5bf47483352c0d9c | 0f0a3c2851134d49955f3c85afbb0b1bb47c3e07 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-10T17:31:40+02:00 | GitHub | noreply@github.com | 2025-09-10T17:31:40+02:00 | | ggml-cpu : fix padding in ggml_timestep_embedding (#15917) |
| 646 | 0f0a3c2851134d49955f3c85afbb0b1bb47c3e07 | 33daece86b65607451d0d4378d2d04ba6a20ad55 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-10T17:52:35+03:00 | GitHub | noreply@github.com | 2025-09-10T17:52:35+03:00 | | metal : make the backend async (#15906) |
| 647 | 33daece86b65607451d0d4378d2d04ba6a20ad55 | e7b6d83b524bbc24a1343d862de6dba8e8eddbd6 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-10T15:39:57+02:00 | GitHub | noreply@github.com | 2025-09-10T15:39:57+02:00 | | ci : add caching for ROCm installation in release workflow (#15924) |
| 648 | e7b6d83b524bbc24a1343d862de6dba8e8eddbd6 | 2cfef4d117d67ab1dec002915b48a15d11ee1973 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-10T14:17:09+02:00 | GitHub | noreply@github.com | 2025-09-10T14:17:09+02:00 | | tests : filter out no-ops from coverage report (#15900) |
| 649 | 2cfef4d117d67ab1dec002915b48a15d11ee1973 | 09e72a037c77b6823fde9e1e4cf28e815b9c41ab | j-k | dev@j-k.io | 2025-09-10T12:51:28+01:00 | GitHub | noreply@github.com | 2025-09-10T14:51:28+03:00 | | media : add transparent icon svg and png [no ci] (#15891) |
| 650 | 09e72a037c77b6823fde9e1e4cf28e815b9c41ab | 10d8b2b6b0ac2cae252a80b4daea5da55ab63c2f | Jesse | jesse@createthis.com | 2025-09-10T07:28:47-04:00 | GitHub | noreply@github.com | 2025-09-10T14:28:47+03:00 | | gitignore : Ignore vim swap files in tests (#15901) |
| 651 | 10d8b2b6b0ac2cae252a80b4daea5da55ab63c2f | 28b5f190ef1dbea5edf82dbc8b4407b721fadd13 | Chenguang Li | 757486878@qq.com | 2025-09-10T18:42:00+08:00 | GitHub | noreply@github.com | 2025-09-10T18:42:00+08:00 | | CANN: Add ROPE sin/cos cache for reuse (#15912) |
| 652 | 28b5f190ef1dbea5edf82dbc8b4407b721fadd13 | 86587da03bd78df8f4e7d8b111a0c1d2494d6ed0 | Chenguang Li | 757486878@qq.com | 2025-09-10T15:29:12+08:00 | GitHub | noreply@github.com | 2025-09-10T15:29:12+08:00 | | CANN: implement LRU cache for ACL graphs (#15814) |
| 653 | 86587da03bd78df8f4e7d8b111a0c1d2494d6ed0 | ff02caf9eed261423289d1531a56536fbf57bfc2 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-10T05:33:58+02:00 | GitHub | noreply@github.com | 2025-09-10T05:33:58+02:00 | | llama : check returned fn ptrs from ggml_backend_reg_get_proc_address (#15893) |
| 654 | ff02caf9eed261423289d1531a56536fbf57bfc2 | ae355f6f7108540297d8b7f7ae71d20fe610a0b7 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-10T05:23:19+02:00 | GitHub | noreply@github.com | 2025-09-10T05:23:19+02:00 | | ci : cache ROCm installation in windows-latest-cmake-hip (#15887) |
| 655 | ae355f6f7108540297d8b7f7ae71d20fe610a0b7 | 4f63cd705c7b6f457f36c63fdc053e07f6f3cc6b | Ruben Ortlam | picard12@live.de | 2025-09-09T22:26:03+02:00 | GitHub | noreply@github.com | 2025-09-09T22:26:03+02:00 | | vulkan: throw the oom error instead of no memory type found (#15905) |
| 656 | 4f63cd705c7b6f457f36c63fdc053e07f6f3cc6b | 17bc5a815f0bebc1844c99796760cd7df5e9b8d9 | Jeff Bolz | jbolz@nvidia.com | 2025-09-09T07:41:15-05:00 | GitHub | noreply@github.com | 2025-09-09T14:41:15+02:00 | | vulkan: Fix OOB accesses in soft_max_back (#15861) |
| 657 | 17bc5a815f0bebc1844c99796760cd7df5e9b8d9 | ed54e32558ffe0f2c3b1d31f3227211fae05bd49 | Johannes Gäßler | johannesg@5d6.de | 2025-09-09T14:04:43+02:00 | GitHub | noreply@github.com | 2025-09-09T14:04:43+02:00 | | HIP: use v_dot2_f32_f16 instruction for FA (#15884) |
| 658 | ed54e32558ffe0f2c3b1d31f3227211fae05bd49 | a972faebed5fdc4a3d2a844d92d476058c02e02d | lksj92hs | 134250687+lksj92hs@users.noreply.github.com | 2025-09-09T15:01:15+03:00 | GitHub | noreply@github.com | 2025-09-09T14:01:15+02:00 | | Workaround for subgroup arithmetic failing on MoltenVK with AMD GPUs (issue 15846) (#15886) |
| 659 | a972faebed5fdc4a3d2a844d92d476058c02e02d | 550cf726e133fd0a069d991287fd3a2a3e3e1cbd | Aman Gupta | amangupta052@gmail.com | 2025-09-09T14:38:02+08:00 | GitHub | noreply@github.com | 2025-09-09T14:38:02+08:00 | | CUDA: Add mul_mat_id support for the mmf kernel (#15767) |
| 660 | 550cf726e133fd0a069d991287fd3a2a3e3e1cbd | c252ce67c4b99f056eeb18d42a289e28fee03475 | Johannes Gäßler | johannesg@5d6.de | 2025-09-09T08:11:01+02:00 | GitHub | noreply@github.com | 2025-09-09T08:11:01+02:00 | | CUDA: fix GET_ROWS for large tensors (#15882) |
| 661 | c252ce67c4b99f056eeb18d42a289e28fee03475 | 70cd37dbbebdb7a2f84f08207f18eabb0b291a55 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-09T08:42:10+03:00 | GitHub | noreply@github.com | 2025-09-09T08:42:10+03:00 | | contrib : add notes about merging PRs (#15881) |
| 662 | 70cd37dbbebdb7a2f84f08207f18eabb0b291a55 | acc1b008cfd95e63d5f99a370dbffb98e5a99d2c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-09T06:06:52+02:00 | GitHub | noreply@github.com | 2025-09-09T06:06:52+02:00 | | requirements : update transformers/torch for Embedding Gemma (#15828) |
| 663 | acc1b008cfd95e63d5f99a370dbffb98e5a99d2c | 7057faf64b514e991e2f70147f82bb13d544b1c0 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-09-09T06:05:55+02:00 | GitHub | noreply@github.com | 2025-09-09T06:05:55+02:00 | | model-conversion : add extra debugging support for model conversion (#15877) |
| 664 | 7057faf64b514e991e2f70147f82bb13d544b1c0 | fe1c92cd7bb491e9c5767c9de413f2b7dc1e832b | Aldehir Rojas | hello@alde.dev | 2025-09-08T16:14:32-05:00 | GitHub | noreply@github.com | 2025-09-08T16:14:32-05:00 | | json : support `enum` values within `allOf` (#15830) |
| 665 | fe1c92cd7bb491e9c5767c9de413f2b7dc1e832b | e68aa10d8f3d26fdad5b912540362d79de5460e3 | j-k | dev@j-k.io | 2025-09-08T19:57:01+01:00 | GitHub | noreply@github.com | 2025-09-08T21:57:01+03:00 | | media : add llama1 icon (#15878) |
| 666 | e68aa10d8f3d26fdad5b912540362d79de5460e3 | 0a16bf52e6874369cb4fd2bdc4863ef984b688e7 | Jeff Bolz | jbolz@nvidia.com | 2025-09-08T13:10:07-05:00 | GitHub | noreply@github.com | 2025-09-09T02:10:07+08:00 | | vulkan: sort graph to allow more parallel execution (#15850) |
| 667 | 0a16bf52e6874369cb4fd2bdc4863ef984b688e7 | 88021565f08e0b7c4e07ac089a15ec16fae9166c | Aman Gupta | amangupta052@gmail.com | 2025-09-09T01:23:46+08:00 | GitHub | noreply@github.com | 2025-09-09T01:23:46+08:00 | | CUDA: generate_cu_files.py - add missing mxfp4 (#15880) |
| 668 | 88021565f08e0b7c4e07ac089a15ec16fae9166c | 56920f56651908d5cce7a310dabf54ac4f6fbb7f | Jesse | jesse@createthis.com | 2025-09-08T10:59:48-04:00 | GitHub | noreply@github.com | 2025-09-08T16:59:48+02:00 | | chat : Deepseek V3.1 reasoning and tool calling support (OpenAI Style) (#15533) |
| 669 | 56920f56651908d5cce7a310dabf54ac4f6fbb7f | b0d52998b962bd2681c34bf52af993af79f178b8 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-08T21:50:05+07:00 | GitHub | noreply@github.com | 2025-09-08T16:50:05+02:00 | | server : bring back timings_per_token (#15879) |
| 670 | b0d52998b962bd2681c34bf52af993af79f178b8 | f28d4f4ac963f182ea9d0fe9e269f3f5f3782aaf | Georgi Gerganov | ggerganov@gmail.com | 2025-09-08T13:56:51+03:00 | GitHub | noreply@github.com | 2025-09-08T13:56:51+03:00 | | cuda : fix supports_op condition for get_rows when number of blocks is too large (#15868) |
| 671 | f28d4f4ac963f182ea9d0fe9e269f3f5f3782aaf | 9fcb29f22f5c33c04c7f0daebb24057899d67a1a | Georgi Gerganov | ggerganov@gmail.com | 2025-09-08T13:34:56+03:00 | GitHub | noreply@github.com | 2025-09-08T13:34:56+03:00 | | metal : refactor + optimize (#15857) |
| 672 | 9fcb29f22f5c33c04c7f0daebb24057899d67a1a | 5ef22d281de9c5eaaf616874bc490b89241128cb | Xuan-Son Nguyen | son@huggingface.co | 2025-09-08T17:33:01+07:00 | GitHub | noreply@github.com | 2025-09-08T12:33:01+02:00 | | ggml: allow casting between f32 and i32 (#15783) |
| 673 | 5ef22d281de9c5eaaf616874bc490b89241128cb | 233d773d02c37982badddf3994e9953a175f34da | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-08T11:55:44+02:00 | GitHub | noreply@github.com | 2025-09-08T12:55:44+03:00 | | CUDA: non-contiguous src0 not supported for PAD (#15869) |
| 674 | 233d773d02c37982badddf3994e9953a175f34da | a885dcff11a7b73f9377812d6151f6b15d307de0 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-08T09:44:34+02:00 | GitHub | noreply@github.com | 2025-09-08T09:44:34+02:00 | | convert : force setting sliding_window from original config (#15867) |
| 675 | a885dcff11a7b73f9377812d6151f6b15d307de0 | 663027fd5490438ce9c27ea866a560e1e268d11f | Georgi Gerganov | ggerganov@gmail.com | 2025-09-08T10:27:07+03:00 | GitHub | noreply@github.com | 2025-09-08T10:27:07+03:00 | | batched-bench : fix llama_synchronize usage during prompt processing (#15835) |
| 676 | 663027fd5490438ce9c27ea866a560e1e268d11f | cf0e3ba1500bd23635b444f64ba23cbdb56c92ef | Georgi Gerganov | ggerganov@gmail.com | 2025-09-08T10:26:36+03:00 | GitHub | noreply@github.com | 2025-09-08T10:26:36+03:00 | | context : fix n_outputs during reserve (#15858) |
| 677 | cf0e3ba1500bd23635b444f64ba23cbdb56c92ef | d413dca00360a7e4cb71441dacecfa32556fcc31 | Georgi Gerganov | ggerganov@gmail.com | 2025-09-08T10:25:33+03:00 | GitHub | noreply@github.com | 2025-09-08T10:25:33+03:00 | | model : avoid ggml_cont_3d for fused QKV weights (#15662) |
| 678 | d413dca00360a7e4cb71441dacecfa32556fcc31 | 85ca66a74676e6d5df4433016488e039a4b464ae | Jeff Bolz | jbolz@nvidia.com | 2025-09-07T23:23:41-05:00 | GitHub | noreply@github.com | 2025-09-07T23:23:41-05:00 | | tests: large sizes for get_rows (#15687) |
| 679 | 85ca66a74676e6d5df4433016488e039a4b464ae | 3976dfbe00f02a62c0deca32c46138e4f0ca81d8 | Chenguang Li | 757486878@qq.com | 2025-09-08T10:03:29+08:00 | GitHub | noreply@github.com | 2025-09-08T10:03:29+08:00 | | CANN: Stream sync between devices for acl_graph (#15809) |
| 680 | 3976dfbe00f02a62c0deca32c46138e4f0ca81d8 | d36e61c580bf7fc7879c443c542312a42b718e11 | Jeff Bolz | jbolz@nvidia.com | 2025-09-07T13:50:26-05:00 | GitHub | noreply@github.com | 2025-09-07T13:50:26-05:00 | | vulkan: support im2col_3d (#15795) |
| 681 | d36e61c580bf7fc7879c443c542312a42b718e11 | c97b5e5854b47b18a248d77edb693c63018a0865 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-08T02:18:28+08:00 | GitHub | noreply@github.com | 2025-09-08T02:18:28+08:00 | | ggml-cpu: clean up s390x SIMD (#15855) |
| 682 | c97b5e5854b47b18a248d77edb693c63018a0865 | 267e99867f09bec8bcc2e424ad9bcddd6cccf9d0 | Jeff Bolz | jbolz@nvidia.com | 2025-09-07T12:00:49-05:00 | GitHub | noreply@github.com | 2025-09-07T19:00:49+02:00 | | vulkan: Support pad_ext (#15794) |
| 683 | 267e99867f09bec8bcc2e424ad9bcddd6cccf9d0 | 3b15924d71237a43bb5ad71f5b885ee66a821342 | Jeff Bolz | jbolz@nvidia.com | 2025-09-07T11:53:07-05:00 | GitHub | noreply@github.com | 2025-09-07T18:53:07+02:00 | | vulkan: Use larger loads in scalar/coopmat1 matmul (#15729) |
| 684 | 3b15924d71237a43bb5ad71f5b885ee66a821342 | 79bc429262268ad2ac8a364cfe6c2d6b9c5f008a | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-07T10:19:45+02:00 | GitHub | noreply@github.com | 2025-09-07T11:19:45+03:00 | | ggml WebGPU: remove userdata from request adapter callback (#15527) |
| 685 | 79bc429262268ad2ac8a364cfe6c2d6b9c5f008a | c4df49a42d396bdf7344501813e7de53bc9e7bb3 | Johannes Gäßler | johannesg@5d6.de | 2025-09-07T00:26:28+02:00 | GitHub | noreply@github.com | 2025-09-07T00:26:28+02:00 | | CUDA: faster tile FA (Pascal/AMD), headsize 256 (#15769) |
| 686 | c4df49a42d396bdf7344501813e7de53bc9e7bb3 | 3c3635d2f20424d557b5b0605a2a356214ffe048 | Charles Xu | charles.xu@arm.com | 2025-09-06T16:08:43+02:00 | GitHub | noreply@github.com | 2025-09-06T22:08:43+08:00 | | kleidiai: generalize compute_forward_kv_cache to compute_forward_fp16 (#15817) |
| 687 | 3c3635d2f20424d557b5b0605a2a356214ffe048 | 61bdfd5298a78593be649a1035ee2a120b13c4f0 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-06T19:45:24+07:00 | GitHub | noreply@github.com | 2025-09-06T14:45:24+02:00 | | server : speed up tests (#15836) |
| 688 | 61bdfd5298a78593be649a1035ee2a120b13c4f0 | 01806e77714ae8a78130d432945b959a0956c56f | Xuan-Son Nguyen | son@huggingface.co | 2025-09-06T18:35:04+07:00 | GitHub | noreply@github.com | 2025-09-06T13:35:04+02:00 | | server : implement prompt processing progress report in stream mode (#15827) |
| 689 | 186415d59552bd5c90549cd8e9be3cc287621fd9 | fd621880f3fd908424e675a41715a2dc760247a2 | Aaron Teo | aaron.teo1@ibm.com | 2025-09-06T11:27:28+08:00 | GitHub | noreply@github.com | 2025-09-06T11:27:28+08:00 | | ggml-cpu: drop support for nnpa intrinsics (#15821) |
| 690 | fd621880f3fd908424e675a41715a2dc760247a2 | 4281c7b315f8a904549fe527039b976e08098d1a | Gabe Goodhart | ghart@us.ibm.com | 2025-09-05T17:32:39-06:00 | GitHub | noreply@github.com | 2025-09-05T17:32:39-06:00 | | aLoRA Support (#15327) |
| 691 | 4281c7b315f8a904549fe527039b976e08098d1a | 5fac79cbc77b6d12c9feb5f34fc63586b35fd561 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-09-06T01:21:15+02:00 | GitHub | noreply@github.com | 2025-09-06T01:21:15+02:00 | | ci : exempt correct research label (#15825) |
| 692 | 5fac79cbc77b6d12c9feb5f34fc63586b35fd561 | 408ff524b40baf4f51a81d42a9828200dd4fcb6b | Gabe Goodhart | ghart@us.ibm.com | 2025-09-05T14:31:24-06:00 | GitHub | noreply@github.com | 2025-09-05T14:31:24-06:00 | | Thinking model disabled assistant prefill (#15404) |
| 693 | 408ff524b40baf4f51a81d42a9828200dd4fcb6b | 5143fa895e7725c5bd2135daf7d8f793d98fa91c | Eric Curtin | ericcurtin17@gmail.com | 2025-09-05T19:43:59+01:00 | GitHub | noreply@github.com | 2025-09-05T19:43:59+01:00 | | Implement --log-colors with always/never/auto (#15792) |
| 694 | 5143fa895e7725c5bd2135daf7d8f793d98fa91c | 3a550b5ca4565c9e28f63880d47840feb27d0ff6 | Johannes Gäßler | johannesg@5d6.de | 2025-09-05T16:07:02+02:00 | GitHub | noreply@github.com | 2025-09-05T16:07:02+02:00 | | CUDA: fastdiv, launch bounds for mmvq + q8_1 quant (#15802) |
| 695 | 3a550b5ca4565c9e28f63880d47840feb27d0ff6 | a81283820a466f2ace06ce4d4bc9512761f9365f | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-05T14:49:21+02:00 | GitHub | noreply@github.com | 2025-09-05T13:49:21+01:00 | | tests : add --list-ops and --show-coverage options (#15745) |
| 696 | a81283820a466f2ace06ce4d4bc9512761f9365f | c610b6c11b1ef7d678671dcf15acd7187a7ad8f3 | Erik Scholz | Green-Sky@users.noreply.github.com | 2025-09-05T11:34:28+02:00 | GitHub | noreply@github.com | 2025-09-05T11:34:28+02:00 | | gguf: gguf_writer refactor (#15691) |
| 697 | c610b6c11b1ef7d678671dcf15acd7187a7ad8f3 | 5d6688de08e73acc2532d668380801ed79d704eb | Georgi Gerganov | ggerganov@gmail.com | 2025-09-05T10:39:22+03:00 | GitHub | noreply@github.com | 2025-09-05T10:39:22+03:00 | | kv-cache : fix SWA checks + disable cacheless iSWA (#15811) |
| 698 | dcbe998a6d8f13b6c9cb36bbe7b355ced2ac3842 | 87932ec016f0ec81e98a13ce5212851f8c12458a | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-05T14:34:07+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-09-05T14:34:07+08:00 | | add calrt API used in server side |
| 699 | 5d6688de08e73acc2532d668380801ed79d704eb | 4fd1242bef6cb2325b4ff1c1a80f3b54b64508a6 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-05T04:36:23+02:00 | GitHub | noreply@github.com | 2025-09-05T04:36:23+02:00 | | model-conversion : add --embeddings flag to modelcard.template [no ci] (#15801) |
| 700 | 4fd1242bef6cb2325b4ff1c1a80f3b54b64508a6 | b2426e469e2fdb6c44216d56baa4cfff4f39ae00 | ExtReMLapin | 3909752+ExtReMLapin@users.noreply.github.com | 2025-09-05T01:24:08+02:00 | GitHub | noreply@github.com | 2025-09-05T01:24:08+02:00 | | chat : fixed crash when Hermes 2 <tool_call> had a newline before it (#15639) |
| 701 | b2426e469e2fdb6c44216d56baa4cfff4f39ae00 | 9e2b1e83c68a38ea0c64f726dd979439bd02189b | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-09-05T01:22:22+02:00 | GitHub | noreply@github.com | 2025-09-05T01:22:22+02:00 | | chat : nemotron thinking & toolcalling support (#15676) |
| 702 | 9e2b1e83c68a38ea0c64f726dd979439bd02189b | fb15d649ed14ab447eeab911e0c9d21e35fb243e | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-09-05T01:05:12+02:00 | GitHub | noreply@github.com | 2025-09-05T01:05:12+02:00 | | scripts : add Jinja tester PySide6 simple app (#15756) |
| 703 | fb15d649ed14ab447eeab911e0c9d21e35fb243e | 856ed0947f27b4ec3ad269fceda0402fbab263d3 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-04T18:10:29+02:00 | GitHub | noreply@github.com | 2025-09-04T18:10:29+02:00 | | llama : add support for EmbeddingGemma 300m (#15798) |
| 704 | 856ed0947f27b4ec3ad269fceda0402fbab263d3 | d1e2adba65208d6de3d7f6f23b00c562ff5bb777 | Gabe Goodhart | ghart@us.ibm.com | 2025-09-04T09:53:22-06:00 | GitHub | noreply@github.com | 2025-09-04T18:53:22+03:00 | | metal : Add template specialization for mul_mm_id w/ ne20 == 10 (#15799) |
| 705 | d1e2adba65208d6de3d7f6f23b00c562ff5bb777 | c1c354e44c06d259679bb5bb7a8fa9f0b28480e4 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-04T15:40:44+02:00 | GitHub | noreply@github.com | 2025-09-04T15:40:44+02:00 | | llama : set n_outputs to 1 to avoid 0 outputs mean-pooling (#15791) |
| 706 | c1c354e44c06d259679bb5bb7a8fa9f0b28480e4 | a68d9144262f1d0ef4f6ba7ad4a7e73e977ba78c | Chenguang Li | 757486878@qq.com | 2025-09-04T20:20:14+08:00 | GitHub | noreply@github.com | 2025-09-04T20:20:14+08:00 | | CANN: Refactor ND to NZ workspace to be per-device (#15763) |
| 707 | a68d9144262f1d0ef4f6ba7ad4a7e73e977ba78c | badb80cadbc40e047b30c43611aba575fc8d6845 | Xuan-Son Nguyen | son@huggingface.co | 2025-09-04T11:50:23+02:00 | GitHub | noreply@github.com | 2025-09-04T11:50:23+02:00 | | server: add exceed_context_size_error type (#15780) |
| 708 | badb80cadbc40e047b30c43611aba575fc8d6845 | 0a1b3982cd0bd18730d50a693053b88c13fd04a6 | Eric Curtin | ecurtin@redhat.com | 2025-09-04T10:49:44+01:00 | GitHub | noreply@github.com | 2025-09-04T10:49:44+01:00 | | Document the new max GPU layers default in help (#15771) |
| 709 | 0a1b3982cd0bd18730d50a693053b88c13fd04a6 | 5421f63ab08ca1e1a093662a5ccd0117e461185f | leejet | leejet714@gmail.com | 2025-09-04T16:38:49+08:00 | GitHub | noreply@github.com | 2025-09-04T10:38:49+02:00 | | ggml: add ops for WAN video model (cuda && cpu) (#15669) |
| 710 | 5421f63ab08ca1e1a093662a5ccd0117e461185f | 820bc9853100708011036e10f40b923968dd9b66 | hipudding | huafengchun@gmail.com | 2025-09-04T15:12:30+08:00 | GitHub | noreply@github.com | 2025-09-04T15:12:30+08:00 | | CANN: Fix precision issue on 310I DUO multi-devices (#15784) |
| 711 | 820bc9853100708011036e10f40b923968dd9b66 | 239b60e8986bbcb944f57311a02ac994431bf652 | rmatif | rmatif@proton.me | 2025-09-04T08:30:28+02:00 | GitHub | noreply@github.com | 2025-09-03T23:30:28-07:00 | | opencl: add hs=40 to FA (#15758) |
| 712 | 239b60e8986bbcb944f57311a02ac994431bf652 | dff7551bfdbdd6e57c13e523d7dcca317640e907 | Chenguang Li | 757486878@qq.com | 2025-09-04T11:03:02+08:00 | GitHub | noreply@github.com | 2025-09-04T11:03:02+08:00 | | CANN: fix acl_rstd allocation size in ggml_cann_rms_norm (#15760) |
| 713 | dff7551bfdbdd6e57c13e523d7dcca317640e907 | 0fce7a1248b74148c1eb0d368b7e18e8bcb96809 | Ruben Ortlam | picard12@live.de | 2025-09-03T22:55:10+02:00 | GitHub | noreply@github.com | 2025-09-03T21:55:10+01:00 | | vulkan: fix mmv subgroup16 selection (#15775) |
| 714 | 0fce7a1248b74148c1eb0d368b7e18e8bcb96809 | 8227695d7a3e5b357cb37fad263f00a8ca6db710 | Jeff Bolz | jbolz@nvidia.com | 2025-09-03T13:33:15-05:00 | GitHub | noreply@github.com | 2025-09-03T20:33:15+02:00 | | vulkan: don't use std::string in load_shaders, to improve compile time (#15724) |
| 715 | 8227695d7a3e5b357cb37fad263f00a8ca6db710 | 0014fb4add3a7d000c14b763538045c6170f57d9 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-03T20:24:50+02:00 | GitHub | noreply@github.com | 2025-09-03T20:24:50+02:00 | | vulkan : update ggml_vk_instance_validation_ext_available (#15666) |
| 716 | 0014fb4add3a7d000c14b763538045c6170f57d9 | 661ae31c9c68201577e70278285b349a5a662caf | Shin-myoung-serp | relent95@naver.com | 2025-09-04T03:22:55+09:00 | GitHub | noreply@github.com | 2025-09-03T20:22:55+02:00 | | ggml vulkan: add hardsigmoid and hardswish operations (#15762) |
| 717 | 661ae31c9c68201577e70278285b349a5a662caf | 407c23786dd0d3a503e9429eead96e611d3950c9 | Oliver Simons | osimons@nvidia.com | 2025-09-03T19:59:16+02:00 | GitHub | noreply@github.com | 2025-09-03T19:59:16+02:00 | | CUDA: Optimize `rms_norm_f32` kernel and its fused variants, giving 1-6% perf E2E (#15715) |
| 718 | 407c23786dd0d3a503e9429eead96e611d3950c9 | cdedb70a998cea7052560fe0b0615a839443564d | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-03T18:28:36+02:00 | GitHub | noreply@github.com | 2025-09-03T18:28:36+02:00 | | model-conversion : fix pyright errors (#15770) |
| 719 | cdedb70a998cea7052560fe0b0615a839443564d | 2c8dac72eb6acd4e20c0da251535dfc46d35178b | Georgi Gerganov | ggerganov@gmail.com | 2025-09-03T18:16:26+03:00 | GitHub | noreply@github.com | 2025-09-03T18:16:26+03:00 | | sampling : optimize dist sampler (#15704) |
| 720 | 2c8dac72eb6acd4e20c0da251535dfc46d35178b | 40a751ea9a94364da73537b86502a808ebe1fc3a | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-03T13:35:49+02:00 | GitHub | noreply@github.com | 2025-09-03T13:35:49+02:00 | | llama : fix incorrect model type for Gemma 270M (#15764) |
| 721 | 40a751ea9a94364da73537b86502a808ebe1fc3a | 5eae9348835037046f33fafd852df7c952c95209 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-03T12:50:47+02:00 | GitHub | noreply@github.com | 2025-09-03T12:50:47+02:00 | | model-conversion : remove hardcoded /bin/bash shebangs [no ci] (#15765) |
| 722 | 5eae9348835037046f33fafd852df7c952c95209 | 05c0380f2ae5c7193445e03ebc7b6de9ce49f1a6 | hipudding | huafengchun@gmail.com | 2025-09-03T16:46:01+08:00 | GitHub | noreply@github.com | 2025-09-03T16:46:01+08:00 | | CANN: Add RoPE contiguous check for 310I DUP device (#15735) |
| 723 | 05c0380f2ae5c7193445e03ebc7b6de9ce49f1a6 | 8c3fdf44ecf08335942f2ba558955f55e88c7991 | xctan | xc-tan@outlook.com | 2025-09-03T16:16:21+08:00 | GitHub | noreply@github.com | 2025-09-03T16:16:21+08:00 | | ggml-cpu : optimize RVV kernels (#15720) |
| 724 | 8c3fdf44ecf08335942f2ba558955f55e88c7991 | f6da8cb86a28f0319b40d9d2a957a26a7d875f8c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-03T09:48:35+02:00 | GitHub | noreply@github.com | 2025-09-03T09:48:35+02:00 | | model-conversion : add missing curl script [no ci] (#15761) |
| 725 | f6da8cb86a28f0319b40d9d2a957a26a7d875f8c | 8a2234ea0c89f212190c176d741b7742f0082582 | hipudding | huafengchun@gmail.com | 2025-09-03T14:08:22+08:00 | GitHub | noreply@github.com | 2025-09-03T14:08:22+08:00 | | CANN: Mask unsupported TRANSPOSE_1D operator (#15733) |
| 726 | 8a2234ea0c89f212190c176d741b7742f0082582 | 3de008208b9b8a33f49f979097a99b4d59e6e521 | Chenguang Li | 757486878@qq.com | 2025-09-03T10:43:53+08:00 | GitHub | noreply@github.com | 2025-09-03T10:43:53+08:00 | | CANN: Fix type float_t to float (#15736) |
| 727 | 3de008208b9b8a33f49f979097a99b4d59e6e521 | 69db8a52e6dd0db81ec05c581150cc1e43b8ac46 | SnA1lGo | 44647694+skrandy@users.noreply.github.com | 2025-09-03T03:27:30+08:00 | GitHub | noreply@github.com | 2025-09-02T21:27:30+02:00 | | fix: resolve unsigned int initialization warning for n_dims/size in gguf.cpp (#15754) |
| 728 | 69db8a52e6dd0db81ec05c581150cc1e43b8ac46 | c466abe1587bde458c806e2134e37750f2edf8fb | Oliver Simons | osimons@nvidia.com | 2025-09-02T19:40:37+02:00 | GitHub | noreply@github.com | 2025-09-03T01:40:37+08:00 | | chore: Update `.clang-format` to use `BinPackArguments=true` (#15744) |
| 729 | c466abe1587bde458c806e2134e37750f2edf8fb | 0a2a3841e8ebc570da343e8cf58a21c1010b41e7 | Johannes Gäßler | johannesg@5d6.de | 2025-09-02T18:17:26+02:00 | GitHub | noreply@github.com | 2025-09-02T18:17:26+02:00 | | llama: -fa 1/0/-1 aliases for -fa on/off/auto (#15746) |
| 730 | 0a2a3841e8ebc570da343e8cf58a21c1010b41e7 | 9961d244f2df6baf40af2f1ddc0927f8d91578c8 | Ruben Ortlam | picard12@live.de | 2025-09-02T16:02:26+02:00 | GitHub | noreply@github.com | 2025-09-02T16:02:26+02:00 | | vulkan: fix shaders gen when no integer dot is available (#15740) |
| 731 | 9961d244f2df6baf40af2f1ddc0927f8d91578c8 | 25f1045f07cf0daf667d63e35618842e3174a8c7 | hipudding | huafengchun@gmail.com | 2025-09-02T17:12:37+08:00 | GitHub | noreply@github.com | 2025-09-02T17:12:37+08:00 | | CANN: Resolve soft_max precision issue (#15730) |
| 732 | 25f1045f07cf0daf667d63e35618842e3174a8c7 | 97669e40735d08db65ef094a8faae8e8411a97db | Jeff Bolz | jbolz@nvidia.com | 2025-09-02T01:37:01-05:00 | GitHub | noreply@github.com | 2025-09-02T14:37:01+08:00 | | vulkan: Fix macro parameter order for f32 matmul shaders (#15716) |
| 733 | 97669e40735d08db65ef094a8faae8e8411a97db | 2f853687b3bce15e143a22f678d1715060fd606c | rmatif | rmatif@proton.me | 2025-09-02T08:26:53+02:00 | GitHub | noreply@github.com | 2025-09-01T23:26:53-07:00 | | opencl: add attn sinks support for FA kernels (#15706) |
| 734 | 2f853687b3bce15e143a22f678d1715060fd606c | ef2af57ddf04ab367f6ba6e8ed100fc56785e4b7 | Chenguang Li | 757486878@qq.com | 2025-09-02T14:07:48+08:00 | GitHub | noreply@github.com | 2025-09-02T14:07:48+08:00 | | CANN: Support eager execution mode under ACL graph compilation (#15712) |
| 735 | ef2af57ddf04ab367f6ba6e8ed100fc56785e4b7 | 5d804a4938d896f9089687a8b99911cdb43b0b94 | hipudding | huafengchun@gmail.com | 2025-09-02T14:05:23+08:00 | GitHub | noreply@github.com | 2025-09-02T14:05:23+08:00 | | CANN: Support ext_factor in rope (#15710) |
| 736 | 5d804a4938d896f9089687a8b99911cdb43b0b94 | d4d8dbe383e8b9600cbe8b42016e3a4529b51219 | Johannes Gäßler | johannesg@5d6.de | 2025-09-02T01:14:55+02:00 | GitHub | noreply@github.com | 2025-09-01T16:14:55-07:00 | | ggml-backend: raise GGML_MAX_SPLIT_INPUTS (#15722) |
| 737 | d4d8dbe383e8b9600cbe8b42016e3a4529b51219 | 35a42edac84690990a75726898d975cd08d220c0 | Gilad S. | 7817232+giladgd@users.noreply.github.com | 2025-09-01T22:17:42+03:00 | GitHub | noreply@github.com | 2025-09-01T21:17:42+02:00 | | vulkan: use memory budget extension to read memory usage (#15545) |
| 738 | 35a42edac84690990a75726898d975cd08d220c0 | fec7911f8febb22c27ce1acb12441800da24cc06 | Jeff Bolz | jbolz@nvidia.com | 2025-09-01T14:01:10-05:00 | GitHub | noreply@github.com | 2025-09-01T21:01:10+02:00 | | vulkan: add missing clamps in new mul_mat_id paths (#15702) |
| 739 | fec7911f8febb22c27ce1acb12441800da24cc06 | 078ce23ea77988f2fe1a42afb18d58bc084d55fa | Ruben Ortlam | picard12@live.de | 2025-09-01T20:58:35+02:00 | GitHub | noreply@github.com | 2025-09-01T20:58:35+02:00 | | vulkan: disable large mmv subgroups on older Nvidia GPUs (#15717) |
| 740 | 078ce23ea77988f2fe1a42afb18d58bc084d55fa | a0c2b207c596d1092a08615de61ab56a7d63515f | s-goto-11 | 206795233+s-goto-11@users.noreply.github.com | 2025-09-02T03:13:49+09:00 | GitHub | noreply@github.com | 2025-09-01T20:13:49+02:00 | | ggml: SVE support for exponential functions (#15145) |
| 741 | a0c2b207c596d1092a08615de61ab56a7d63515f | 4b20d8b7e31718d3fe8b4f12d220a9d19e4f1997 | Prashant Vithule | 119530321+Vithulep@users.noreply.github.com | 2025-09-01T23:43:16+05:30 | GitHub | noreply@github.com | 2025-09-01T20:13:16+02:00 | | ggml: aarch64: Implement SVE F16 kernels for vector functions (#15115) |
| 742 | 4b20d8b7e31718d3fe8b4f12d220a9d19e4f1997 | 02c1813517412f3e00aa6ca7c0273fea64edb492 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-01T23:53:31+08:00 | GitHub | noreply@github.com | 2025-09-01T23:53:31+08:00 | | convert : remove redundant code (#15708) |
| 743 | 02c1813517412f3e00aa6ca7c0273fea64edb492 | 77dee9de97be75b7143a213bc48893e0c0b29af7 | Ruben Ortlam | picard12@live.de | 2025-09-01T16:19:07+02:00 | GitHub | noreply@github.com | 2025-09-01T16:19:07+02:00 | | Vulkan: Add Integer Dot Product mul_mat_vec shader for legacy quants (#14903) |
| 744 | 77dee9de97be75b7143a213bc48893e0c0b29af7 | 4795c91c32fec7165a1364763d4d4f0c93abf933 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-09-01T14:28:49+02:00 | GitHub | noreply@github.com | 2025-09-01T14:28:49+02:00 | | ggml : WebGPU add TRANSPOSE and RESHAPE to supported ops (#15695) |
| 745 | 4795c91c32fec7165a1364763d4d4f0c93abf933 | b66df9d9c942254d03209186ef24ed7c994a576e | Jie Fu (傅杰) | jiefu@tencent.com | 2025-09-01T15:34:59+08:00 | GitHub | noreply@github.com | 2025-09-01T10:34:59+03:00 | | docs : add Hunyuan to models section (#15707) |
| 746 | b66df9d9c942254d03209186ef24ed7c994a576e | b9382c3877c6067feccf182efe9449a2d1cb24c7 | Akarshan Biswas | akarshan@menlo.ai | 2025-09-01T06:55:06+05:30 | GitHub | noreply@github.com | 2025-09-01T06:55:06+05:30 | | CUDA: fix build error from ambiguous __half conversions in conv2d (#15690) |
| 747 | b9382c3877c6067feccf182efe9449a2d1cb24c7 | 3dc7397a2799bdc07bccf637ab7ae5a1e786d1a4 | hipudding | huafengchun@gmail.com | 2025-09-01T08:57:23+08:00 | GitHub | noreply@github.com | 2025-09-01T08:57:23+08:00 | | CANN: Optimize MUL_MAT_ID (#15658) |
| 748 | 3dc7397a2799bdc07bccf637ab7ae5a1e786d1a4 | e92d53b29e393fc4c0f9f1f7c3fe651be8d36faa | hipudding | huafengchun@gmail.com | 2025-09-01T08:57:00+08:00 | GitHub | noreply@github.com | 2025-09-01T08:57:00+08:00 | | CANN: fix RoPE cache issue on multi-device (#15629) |
| 749 | e92d53b29e393fc4c0f9f1f7c3fe651be8d36faa | 0d161f021aa33ec0e90cce96f5d1a88925557327 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-31T20:41:02+03:00 | GitHub | noreply@github.com | 2025-08-31T20:41:02+03:00 | | sampling : optimize samplers by reusing bucket sort (#15665) |
| 750 | 0d161f021aa33ec0e90cce96f5d1a88925557327 | 4efd5a83163ff383285b3a4c2106feabf5c69557 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-31T20:11:58+03:00 | GitHub | noreply@github.com | 2025-08-31T20:11:58+03:00 | | server : enable /slots by default and make it secure (#15630) |
| 751 | 4efd5a83163ff383285b3a4c2106feabf5c69557 | 274966226f87f301ac132da898280ca3142b60e5 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-31T19:43:30+03:00 | GitHub | noreply@github.com | 2025-08-31T19:43:30+03:00 | | metal : fix checks for available FA kernels (#15700) |
| 752 | 274966226f87f301ac132da898280ca3142b60e5 | 9777032dccd67bdc7785aeab7497014a8be8dacc | Diego Devesa | slarengh@gmail.com | 2025-08-31T08:47:05-07:00 | GitHub | noreply@github.com | 2025-08-31T18:47:05+03:00 | | llama : fix fattn reserve call n_seqs parameter (#15699) |
| 753 | 9777032dccd67bdc7785aeab7497014a8be8dacc | 7d3c9f2b217acf0ce5db81ae83d3f375f49ab2c7 | Diego Devesa | slarengh@gmail.com | 2025-08-31T06:49:03-07:00 | GitHub | noreply@github.com | 2025-08-31T15:49:03+02:00 | | llama : separate compute buffer reserve from fattn check (#15696) |
| 754 | 7d3c9f2b217acf0ce5db81ae83d3f375f49ab2c7 | bbbf5ecccb35286521f735239d499eec4279a840 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-31T15:30:20+02:00 | GitHub | noreply@github.com | 2025-08-31T15:30:20+02:00 | | ci : explicitly set fa off or on (#15692) |
| 755 | bbbf5ecccb35286521f735239d499eec4279a840 | c37052ab4d6d1ae73c0e90bc6e560cc6409e1311 | Jeff Bolz | jbolz@nvidia.com | 2025-08-31T03:13:27-05:00 | GitHub | noreply@github.com | 2025-08-31T10:13:27+02:00 | | vulkan: handle large sizes for get_rows (#15686) |
| 756 | c37052ab4d6d1ae73c0e90bc6e560cc6409e1311 | 5c16b9c87d840e4d5d55fa83c732c6b693346f40 | Jeff Bolz | jbolz@nvidia.com | 2025-08-31T02:06:43-05:00 | GitHub | noreply@github.com | 2025-08-31T09:06:43+02:00 | | vulkan: mul_mat_id coopmat2 optimizations (#15546) |
| 757 | 5c16b9c87d840e4d5d55fa83c732c6b693346f40 | b97c9edc59d4a1b4069991aa670411190f4f3a3e | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-31T08:46:42+02:00 | GitHub | noreply@github.com | 2025-08-31T08:46:42+02:00 | | vulkan : remove unused portability_enumeration_ext variable (#15679) |
| 758 | b97c9edc59d4a1b4069991aa670411190f4f3a3e | 94e82c7eadeb8fff0db4bfd1ab6d8cf65fa6f2e0 | Jeff Bolz | jbolz@nvidia.com | 2025-08-31T01:30:54-05:00 | GitHub | noreply@github.com | 2025-08-31T08:30:54+02:00 | | vulkan: Allow fallback to sysmem memory when vidmem is full (#15649) |
| 759 | 94e82c7eadeb8fff0db4bfd1ab6d8cf65fa6f2e0 | 4d74393bcc956ccd7df68a6a06d1a0575cfa712c | Jeff Bolz | jbolz@nvidia.com | 2025-08-31T01:27:57-05:00 | GitHub | noreply@github.com | 2025-08-31T08:27:57+02:00 | | vulkan: clamp matmul and FA results to the max finite value (#15652) |
| 760 | 4d74393bcc956ccd7df68a6a06d1a0575cfa712c | dd892555b0681b7f56d38780f6fdfe00a195160f | Charles Xu | charles.xu@arm.com | 2025-08-30T18:03:42+02:00 | GitHub | noreply@github.com | 2025-08-31T00:03:42+08:00 | | ggml: update kleidiai to v1.13.0 (#15663) |
| 761 | dd892555b0681b7f56d38780f6fdfe00a195160f | e81b8e4b7f5ab870836fad26d154a7507b341b36 | Diego Devesa | slarengh@gmail.com | 2025-08-30T08:51:28-07:00 | GitHub | noreply@github.com | 2025-08-30T23:51:28+08:00 | | Update build.md to remove MSVC arm64 notes (#15684) |
| 762 | e81b8e4b7f5ab870836fad26d154a7507b341b36 | 38ad381f9f5d4dd368a96d844fb19cf501ed9d22 | Johannes Gäßler | johannesg@5d6.de | 2025-08-30T16:32:10+02:00 | GitHub | noreply@github.com | 2025-08-30T16:32:10+02:00 | | llama: use FA + max. GPU layers by default (#15434) |
| 763 | 38ad381f9f5d4dd368a96d844fb19cf501ed9d22 | 696fccf354e9dbdfbce135bc40b44c9dcc64dda9 | Johannes Gäßler | johannesg@5d6.de | 2025-08-30T16:20:32+02:00 | GitHub | noreply@github.com | 2025-08-30T16:20:32+02:00 | | CUDA: use FP32 arithmetic for conv2d (#15683) |
| 764 | 696fccf354e9dbdfbce135bc40b44c9dcc64dda9 | ef476916bba4b44f44be0c98babc1cb025968e75 | Jeff Bolz | jbolz@nvidia.com | 2025-08-30T04:11:22-05:00 | GitHub | noreply@github.com | 2025-08-30T11:11:22+02:00 | | vulkan: Skip syncing for prealloc_y when it is reused (#15544) |
| 765 | ef476916bba4b44f44be0c98babc1cb025968e75 | d82f6aa34a216f5df1945cdfe121ba5e6cd80be0 | Chenguang Li | 757486878@qq.com | 2025-08-30T10:18:35+08:00 | GitHub | noreply@github.com | 2025-08-30T10:18:35+08:00 | | CANN: FIx compiler warnings (#15661) |
| 766 | d82f6aa34a216f5df1945cdfe121ba5e6cd80be0 | 3d16b29c3bb1ec816ac0e782f20d169097063919 | Sergey Alirzaev | l29ah@riseup.net | 2025-08-30T00:12:53+02:00 | GitHub | noreply@github.com | 2025-08-30T00:12:53+02:00 | | server : removed obsolete doc (#15670) |
| 767 | 792b44f2ed9668cce7f267ff0ae4950ed9b4a5de | 81017865ee444cf49ce0136f2be1e41a0270ff91 | ExtReMLapin | 3909752+ExtReMLapin@users.noreply.github.com | 2025-08-29T19:25:40+02:00 | GitHub | noreply@github.com | 2025-08-29T20:25:40+03:00 | | server : add documentation for `parallel_tool_calls` param (#15647) |
| 768 | 81017865ee444cf49ce0136f2be1e41a0270ff91 | 60e5eee31f1af9bb579ac45380e3857d610020b9 | Aman Gupta | amangupta052@gmail.com | 2025-08-29T21:30:06+08:00 | GitHub | noreply@github.com | 2025-08-29T21:30:06+08:00 | | CUDA: fix bug in rms_norm fusion (#15660) |
| 769 | 60e5eee31f1af9bb579ac45380e3857d610020b9 | 009b709d6efd24820ac67765ed339a72dc797814 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-08-29T14:53:41+02:00 | GitHub | noreply@github.com | 2025-08-29T14:53:41+02:00 | | chat : Seed OSS thinking + tool call support (#15552) |
| 770 | 009b709d6efd24820ac67765ed339a72dc797814 | e8d99dd0b67f2ecc1e45fca8074a3a18c3e036d2 | Aman Gupta | amangupta052@gmail.com | 2025-08-29T11:35:58+08:00 | GitHub | noreply@github.com | 2025-08-29T11:35:58+08:00 | | CUDA: fuse adds, fuse add with rms norm (#15631) |
| 771 | e8d99dd0b67f2ecc1e45fca8074a3a18c3e036d2 | a8bca68f727844e7dcf24a956003b3c2039ea563 | Gabe Goodhart | ghart@us.ibm.com | 2025-08-28T18:39:31-06:00 | GitHub | noreply@github.com | 2025-08-28T18:39:31-06:00 | | nvidia nemotron nano v2 (nemotronh) (#15507) |
| 772 | a8bca68f727844e7dcf24a956003b3c2039ea563 | c97dc093912ad014f6d22743ede0d4d7fd82365a | Gabe Goodhart | ghart@us.ibm.com | 2025-08-28T15:27:36-05:00 | GitHub | noreply@github.com | 2025-08-28T15:27:36-05:00 | | fix: Compute the full sum in llama-eval-callback, not just the sum of printed values (#15637) |
| 773 | c97dc093912ad014f6d22743ede0d4d7fd82365a | 6c442f42ff25564a0cd6b1435d9abc1b0178eac5 | mnehete32 | 33429707+mnehete32@users.noreply.github.com | 2025-08-29T00:03:03+05:30 | GitHub | noreply@github.com | 2025-08-28T20:33:03+02:00 | | CUDA: add conv2d (#15635) |
| 774 | 6c442f42ff25564a0cd6b1435d9abc1b0178eac5 | 73804145ab6c06888e314046abf08f66a0133679 | Aaron Teo | aaron.teo1@ibm.com | 2025-08-28T22:39:27+08:00 | GitHub | noreply@github.com | 2025-08-28T22:39:27+08:00 | | ggml-cpu: fix invalid hsum build in debug s390x (#15634) |
| 775 | 73804145ab6c06888e314046abf08f66a0133679 | c8d0d14e77c3c45df5cbbddde9c2b4c377ec4e7a | compilade | git@compilade.net | 2025-08-28T10:11:36-04:00 | GitHub | noreply@github.com | 2025-08-28T10:11:36-04:00 | | ggml : fix SSM_SCAN for n_groups > 1 (#15625) |
| 776 | c8d0d14e77c3c45df5cbbddde9c2b4c377ec4e7a | 84ab83cc0b4b7e769451ee48e4c7d1acef91ef25 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-28T17:09:05+03:00 | GitHub | noreply@github.com | 2025-08-28T17:09:05+03:00 | | kv-cache : fix find_slot to not search for continuous slot (#15638) |
| 777 | 84ab83cc0b4b7e769451ee48e4c7d1acef91ef25 | 55042b3692cb1467c9ee15c62c4a9fbf180f89e3 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-28T15:49:50+02:00 | GitHub | noreply@github.com | 2025-08-28T15:49:50+02:00 | | model : jina-embeddings-v3 support (#13693) |
| 778 | 55042b3692cb1467c9ee15c62c4a9fbf180f89e3 | 8a4280ce431da6b33e5a95ae1fd61472c8c3f8cc | Aman Gupta | amangupta052@gmail.com | 2025-08-28T19:23:22+08:00 | GitHub | noreply@github.com | 2025-08-28T19:23:22+08:00 | | scripts: add sqlite3 check for compare-commits.sh (#15633) |
| 779 | 8a4280ce431da6b33e5a95ae1fd61472c8c3f8cc | 64387f6e95434b393ac3df285864692b7fd9c4d2 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-28T12:27:02+03:00 | GitHub | noreply@github.com | 2025-08-28T12:27:02+03:00 | | kv-cache : remove LLAMA_SET_ROWS checks (#15505) |
| 780 | 64387f6e95434b393ac3df285864692b7fd9c4d2 | d35a1e8c41f747548775225973a99507896a8c61 | Aleksei Nikiforov | 103434461+AlekseiNikiforovIBM@users.noreply.github.com | 2025-08-28T10:56:41+02:00 | GitHub | noreply@github.com | 2025-08-28T16:56:41+08:00 | | gguf-py: byteswapping improvements (#12851) |
| 781 | d35a1e8c41f747548775225973a99507896a8c61 | 46d9caa27a0281150e8cf082308c0f9e7576ebe5 | Joshua Cogliati | jrincayc@users.noreply.github.com | 2025-08-28T01:48:20-06:00 | GitHub | noreply@github.com | 2025-08-28T10:48:20+03:00 | | cli : change log to warning to explain reason for stopping (#15604) |
| 782 | 46d9caa27a0281150e8cf082308c0f9e7576ebe5 | 5a0e3ef6f00c658fbae53797f02d5a360ebf8fec | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-28T09:26:48+02:00 | GitHub | noreply@github.com | 2025-08-28T09:26:48+02:00 | | model-conversion : add mmproj conversion target (#15628) |
| 783 | 5a0e3ef6f00c658fbae53797f02d5a360ebf8fec | fbef0fad7a7c765939f6c9e322fa05cd52cf0c15 | matiaslin | 45382001+matiaslin@users.noreply.github.com | 2025-08-27T17:32:36-07:00 | GitHub | noreply@github.com | 2025-08-28T02:32:36+02:00 | | cuda: Add cublasLt_static linking when GGML_STATIC is enabled (#15622) |
| 784 | fbef0fad7a7c765939f6c9e322fa05cd52cf0c15 | da54f9f1a2db07aaae390024ac466e7867685d94 | Johannes Gäßler | johannesg@5d6.de | 2025-08-27T20:58:09+02:00 | GitHub | noreply@github.com | 2025-08-27T20:58:09+02:00 | | server: higher timeout for tests (#15621) |
| 785 | da54f9f1a2db07aaae390024ac466e7867685d94 | 47373271f971aa5a0a6462b286f4a7b5bd4ba644 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-27T15:48:07+03:00 | GitHub | noreply@github.com | 2025-08-27T15:48:07+03:00 | | presets : add qwen3-30B-a3b FIM (#15616) |
| 786 | 47373271f971aa5a0a6462b286f4a7b5bd4ba644 | 1bded5a3b3376b2aa3ba2f11ade5910c550f354b | uvos | carl@uvos.xyz | 2025-08-27T13:58:54+02:00 | GitHub | noreply@github.com | 2025-08-27T13:58:54+02:00 | | HIP: Enable support for ggml_backend_cuda_register_host_buffer (#15615) |
| 787 | 1bded5a3b3376b2aa3ba2f11ade5910c550f354b | 1e7489745a74996fc36e8fd05b73aa16bc184e0c | Georgi Gerganov | ggerganov@gmail.com | 2025-08-27T13:55:12+03:00 | GitHub | noreply@github.com | 2025-08-27T13:55:12+03:00 | | kv-cache : better estimate of n_kv for multi-sequence batches (#15610) |
| 788 | 1e7489745a74996fc36e8fd05b73aa16bc184e0c | 1cf123a343ab7ca5586aacb9e0a1d2de7fe33be4 | Chenguang Li | 757486878@qq.com | 2025-08-27T17:21:41+08:00 | GitHub | noreply@github.com | 2025-08-27T17:21:41+08:00 | | CANN: refactor mask handling and improve performance in FA (#15561) |
| 789 | 1cf123a343ab7ca5586aacb9e0a1d2de7fe33be4 | fcca2182a18c786c57d02d1d1927204b133f1fdc | xctan | xc-tan@outlook.com | 2025-08-27T16:44:22+08:00 | GitHub | noreply@github.com | 2025-08-27T16:44:22+08:00 | | ggml-cpu : add basic RVV support for vector f32 ops (#15057) |
| 790 | fcca2182a18c786c57d02d1d1927204b133f1fdc | 86076f92de1a547ef87c304facca3b4b9fae6c21 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-27T10:28:53+02:00 | GitHub | noreply@github.com | 2025-08-27T10:28:53+02:00 | | common : add -m to bash completion for --model [no ci] (#15591) |
| 791 | 86076f92de1a547ef87c304facca3b4b9fae6c21 | bcbddcd54f0d5c22eab180831fdea6484107112f | rmatif | rmatif@proton.me | 2025-08-27T08:36:05+02:00 | GitHub | noreply@github.com | 2025-08-26T23:36:05-07:00 | | OpenCL: add fused group_norm/norm, mul, add (#15314) |
| 792 | 87932ec016f0ec81e98a13ce5212851f8c12458a | 0115cfb7c1dec0753aecb34f3af4338bf9de8eeb | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-27T14:29:18+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-27T14:29:18+08:00 | | fix merge problem |
| 793 | bcbddcd54f0d5c22eab180831fdea6484107112f | 8b696861364360770e9f61a3422d32941a477824 | Diego Devesa | slarengh@gmail.com | 2025-08-26T13:14:38-07:00 | GitHub | noreply@github.com | 2025-08-26T22:14:38+02:00 | | tests : fix test-opt with GGML_BACKEND_DL (#15599) |
| 794 | 8b696861364360770e9f61a3422d32941a477824 | 8ce3ff1d91245e158d98d8062cd64b0dd98dcfe3 | Akarshan Biswas | akarshan@menlo.ai | 2025-08-27T00:27:49+05:30 | GitHub | noreply@github.com | 2025-08-27T00:27:49+05:30 | | SYCL: fix rms_norm_mul_add for tensor dim not a multiple of sg_size (#15592) |
| 795 | 8ce3ff1d91245e158d98d8062cd64b0dd98dcfe3 | 44b1efa41acd4df3f56ee0e46f898135ecd1a054 | fidoriel | 49869342+fidoriel@users.noreply.github.com | 2025-08-26T20:05:50+02:00 | GitHub | noreply@github.com | 2025-08-26T20:05:50+02:00 | | mtmd : fix mtmd ios build (#15579) |
| 796 | 44b1efa41acd4df3f56ee0e46f898135ecd1a054 | a6a58d64785cb458ed9de52f391aa38142d38d64 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-08-26T15:42:49Z | GitHub | noreply@github.com | 2025-08-26T15:42:49Z | | tests: add performance test for mul mat id (#15543) |
| 797 | a6a58d64785cb458ed9de52f391aa38142d38d64 | 0373486dbc0dccbdcb3b5fdd65759d88cec06196 | shalinib-ibm | Shalini.Salomi.Bodapati@ibm.com | 2025-08-26T21:05:25+05:30 | GitHub | noreply@github.com | 2025-08-26T23:35:25+08:00 | | llamafile: PowerPC Sgemm Optimization (#15558) |
| 798 | 0373486dbc0dccbdcb3b5fdd65759d88cec06196 | 62cef26ac5b6b7acb635d3dc963813b43952dc2b | Georgi Gerganov | ggerganov@gmail.com | 2025-08-26T17:45:17+03:00 | GitHub | noreply@github.com | 2025-08-26T17:45:17+03:00 | | graph : fix assert in memory-less build_attn (#15590) |
| 799 | 62cef26ac5b6b7acb635d3dc963813b43952dc2b | 8f5afa94c4f929da71f560db7c9f38ef6a783d95 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-26T16:12:29+02:00 | GitHub | noreply@github.com | 2025-08-26T16:12:29+02:00 | | model-conversion : add qat-q4 quantization targets (#15588) |
| 800 | 8f5afa94c4f929da71f560db7c9f38ef6a783d95 | b3964c1e890ef8c947afb36a5124ce6fcb2136d4 | Johannes Gäßler | johannesg@5d6.de | 2025-08-26T16:01:20+02:00 | GitHub | noreply@github.com | 2025-08-26T16:01:20+02:00 | | CUDA: return -1 for nonexistent compiled arch (#15587) |
| 801 | b3964c1e890ef8c947afb36a5124ce6fcb2136d4 | 79a546220c719e6a70627b243a478ab8d84dc9e1 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-26T14:22:14+03:00 | GitHub | noreply@github.com | 2025-08-26T14:22:14+03:00 | | metal : optimize FA vec for large sequences and BS <= 8 (#15566) |
| 802 | 79a546220c719e6a70627b243a478ab8d84dc9e1 | 85cc1ae998e4898d9fa992cb9b8620338cee97bf | Xuan-Son Nguyen | son@huggingface.co | 2025-08-26T12:54:19+02:00 | GitHub | noreply@github.com | 2025-08-26T12:54:19+02:00 | | mtmd : support Kimi VL model (#15458) |
| 803 | 85cc1ae998e4898d9fa992cb9b8620338cee97bf | 1d8d83deaa48d4a5491820d58ff1c0d8cf9d196c | Georgi Gerganov | ggerganov@gmail.com | 2025-08-26T12:47:00+03:00 | GitHub | noreply@github.com | 2025-08-26T12:47:00+03:00 | | context : print graph stats for memory-less contexts (#15586) |
| 804 | 1d8d83deaa48d4a5491820d58ff1c0d8cf9d196c | c4e9239064a564de7b94ee2b401ae907235a8fca | Georgi Gerganov | ggerganov@gmail.com | 2025-08-26T12:46:15+03:00 | GitHub | noreply@github.com | 2025-08-26T12:46:15+03:00 | | metal : improve `MUL_MAT_ID` (#15541) |
| 805 | c4e9239064a564de7b94ee2b401ae907235a8fca | 39842a7f73012eb42816ca4f26411782bd3da7c5 | tc-mb | 157115220+tc-mb@users.noreply.github.com | 2025-08-26T16:05:55+08:00 | GitHub | noreply@github.com | 2025-08-26T10:05:55+02:00 | | model : support MiniCPM-V 4.5 (#15575) |
| 806 | 0115cfb7c1dec0753aecb34f3af4338bf9de8eeb | 3a617cb18121a85eed17ac5585788810c5d9c1e0 0d5a470223fc90b6b6807921d68011ff06ae7f9e | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-26T15:19:56+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-26T15:19:56+08:00 | | Merge branch 'master' into dev |
| 807 | 39842a7f73012eb42816ca4f26411782bd3da7c5 | 0fd90db5858e325358a5fbdaa4e327ee2a79d0d4 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-26T09:08:08+02:00 | GitHub | noreply@github.com | 2025-08-26T09:08:08+02:00 | | gguf-py : remove erroneous FFN_GATE entry (#15583) |
| 808 | 0fd90db5858e325358a5fbdaa4e327ee2a79d0d4 | 4c37636b3ea96f2574eeb7668b93fcc0e64b05dd | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-26T08:51:43+02:00 | GitHub | noreply@github.com | 2025-08-26T09:51:43+03:00 | | metal : remove contiguous assertion for src0 in IM2COL (#15577) |
| 809 | 4c37636b3ea96f2574eeb7668b93fcc0e64b05dd | 34bdbbd7c2b70b848718e95bc48010f6aecd2816 | Yoshi_likes_e4 | 104140648+pt13762104@users.noreply.github.com | 2025-08-26T13:15:33+07:00 | GitHub | noreply@github.com | 2025-08-26T08:15:33+02:00 | | Add a warning for special devices (#15563) |
| 810 | 34bdbbd7c2b70b848718e95bc48010f6aecd2816 | 74f52f77f28a5ad6d6075231afcb8d1ad763ca32 | Jeff Bolz | jbolz@nvidia.com | 2025-08-25T23:42:44-05:00 | GitHub | noreply@github.com | 2025-08-26T06:42:44+02:00 | | vulkan: Remove splitting for mul_mat_id (#15568) |
| 811 | 74f52f77f28a5ad6d6075231afcb8d1ad763ca32 | f7207b0415986dd7f48447149da7de3a82338276 | Qeeweew | 68716978+Qeeweew@users.noreply.github.com | 2025-08-26T05:21:22+08:00 | GitHub | noreply@github.com | 2025-08-25T23:21:22+02:00 | | CUDA: Accelerate MXFP4 table lookup using `__byte_perm` (#15451) |
| 812 | f7207b0415986dd7f48447149da7de3a82338276 | 4d917cd4f64cc37744e76d084659475819fb0728 | lhez | lih@qti.qualcomm.com | 2025-08-25T14:18:09-07:00 | GitHub | noreply@github.com | 2025-08-25T14:18:09-07:00 | | opencl: fix support ops condition for `rms_norm` (#15560) |
| 813 | 4d917cd4f64cc37744e76d084659475819fb0728 | 886b97a5d693550c2da470c091d9d27bf38398f8 | Ruben Ortlam | picard12@live.de | 2025-08-25T17:56:59+02:00 | GitHub | noreply@github.com | 2025-08-25T17:56:59+02:00 | | vulkan: fix min subgroup 16 condition for mmid subgroup optimization (#15565) |
| 814 | 886b97a5d693550c2da470c091d9d27bf38398f8 | 111f8d06f0b9169059779a35f655b343242ccbb6 | Jeff Bolz | jbolz@nvidia.com | 2025-08-25T10:47:16-05:00 | GitHub | noreply@github.com | 2025-08-25T10:47:16-05:00 | | tests: Generate unique input values for count_equal (#15487) |
| 815 | 111f8d06f0b9169059779a35f655b343242ccbb6 | 5eff6ec9b1220b599a43b594b1110487ab6aca08 | Ihar Hrachyshka | ihar.hrachyshka@gmail.com | 2025-08-25T11:27:34-04:00 | GitHub | noreply@github.com | 2025-08-25T18:27:34+03:00 | | metal: fix regression when no metal devices are present (#15531) |
| 816 | 5eff6ec9b1220b599a43b594b1110487ab6aca08 | dfd9b5f6c7586c88588f06a644c131bec071a0a1 | Johannes Gäßler | johannesg@5d6.de | 2025-08-25T17:23:40+02:00 | GitHub | noreply@github.com | 2025-08-25T17:23:40+02:00 | | CUDA: MoE helper in device code, better tile sizes (#15525) |
| 817 | dfd9b5f6c7586c88588f06a644c131bec071a0a1 | 5a6bc6b1a6cb665a944426c2055794950e524bf5 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-25T15:00:43+02:00 | GitHub | noreply@github.com | 2025-08-25T15:00:43+02:00 | | model-conversion : set pooling type to none in logits.cpp (#15564) |
| 818 | 5a6bc6b1a6cb665a944426c2055794950e524bf5 | 6b64f74b55628e4193f4fb00313f07dbd8556528 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-25T14:25:25+02:00 | GitHub | noreply@github.com | 2025-08-25T14:25:25+02:00 | | model-conversion : add model card template for embeddings [no ci] (#15557) |
| 819 | 6b64f74b55628e4193f4fb00313f07dbd8556528 | 0d5a470223fc90b6b6807921d68011ff06ae7f9e | Georgi Gerganov | ggerganov@gmail.com | 2025-08-25T13:56:43+03:00 | GitHub | noreply@github.com | 2025-08-25T13:56:43+03:00 | | batched-bench : fix unified KV cache handling + pp timing (#15562) |
| 820 | 0d5a470223fc90b6b6807921d68011ff06ae7f9e | b0ba31f525a2ad7fb539660be0c8f088c9869f6a | Weizhao Ouyang | o451686892@gmail.com | 2025-08-25T17:15:06+08:00 | GitHub | noreply@github.com | 2025-08-25T11:15:06+02:00 | origin/master, origin/HEAD, master | convert : update Ernie 4.5 dense architecture name (#15555) |
| 821 | b0ba31f525a2ad7fb539660be0c8f088c9869f6a | 7da9fed0d6f1750a8783436f6f313b87e76e6378 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-25T10:14:48+03:00 | GitHub | noreply@github.com | 2025-08-25T10:14:48+03:00 | | metal : add FA kernels for HS=40 (#15559) |
| 822 | 7da9fed0d6f1750a8783436f6f313b87e76e6378 | c247d06f38fc09059c9607a28aa44f5ff6be208d | RunningLeon | mnsheng@yeah.net | 2025-08-25T14:32:16+08:00 | GitHub | noreply@github.com | 2025-08-25T08:32:16+02:00 | | convert : support interns1-mini (#15412) |
| 823 | c247d06f38fc09059c9607a28aa44f5ff6be208d | 043fb27d3808766d8ea8195bbd12359727264402 | Chenguang Li | 757486878@qq.com | 2025-08-25T10:32:21+08:00 | GitHub | noreply@github.com | 2025-08-25T10:32:21+08:00 | | CANN: ROPE cache sin/cos repeat (#15501) |
| 824 | 043fb27d3808766d8ea8195bbd12359727264402 | b730706a49e576fb882dc34d9966345778b3ab0b | Ruben Ortlam | picard12@live.de | 2025-08-24T19:36:36+02:00 | GitHub | noreply@github.com | 2025-08-24T19:36:36+02:00 | | vulkan: apply MUL_MAT_ID subgroup optimization to non-coopmat devices (#15524) |
| 825 | b730706a49e576fb882dc34d9966345778b3ab0b | c9a24fb93208fbbd3da6d903eb75431bfa97e59e | Georgi Gerganov | ggerganov@gmail.com | 2025-08-24T13:07:07+03:00 | GitHub | noreply@github.com | 2025-08-24T13:07:07+03:00 | | kv-cache : support layer reuse (#15504) |
| 826 | c9a24fb93208fbbd3da6d903eb75431bfa97e59e | a9c6ffcbfacee092bfaaa400306fceda18199737 | Jeff Bolz | jbolz@nvidia.com | 2025-08-24T04:24:25-05:00 | GitHub | noreply@github.com | 2025-08-24T11:24:25+02:00 | | vulkan: Support FA with any multiple of 8 head sizes (#15537) |
| 827 | a9c6ffcbfacee092bfaaa400306fceda18199737 | e78cf0d4b1bdbbc2479f11d58ce0c8f51f755875 | Ruben Ortlam | picard12@live.de | 2025-08-24T10:48:53+02:00 | GitHub | noreply@github.com | 2025-08-24T10:48:53+02:00 | | vulkan: enable Conv2D for Apple after MoltenVK fixed the bug (#15526) |
| 828 | e78cf0d4b1bdbbc2479f11d58ce0c8f51f755875 | 710dfc465a68f7443b87d9f792cffba00ed739fe | Jeff Bolz | jbolz@nvidia.com | 2025-08-24T03:48:21-05:00 | GitHub | noreply@github.com | 2025-08-24T10:48:21+02:00 | | vulkan: workaround MoltenVK compile failure in multi_add (#15506) |
| 829 | 710dfc465a68f7443b87d9f792cffba00ed739fe | 611f419cff11e4952228162a1c44cb35dff2274a | Johannes Gäßler | johannesg@5d6.de | 2025-08-23T21:37:06+02:00 | GitHub | noreply@github.com | 2025-08-23T21:37:06+02:00 | | CUDA: fix half2 -> half conversion for HIP (#15529) |
| 830 | 611f419cff11e4952228162a1c44cb35dff2274a | b1afcab804e3281867a5471fbd701e32eb32e512 | Jeff Bolz | jbolz@nvidia.com | 2025-08-23T13:16:17-05:00 | GitHub | noreply@github.com | 2025-08-23T13:16:17-05:00 | | vulkan: optimize rms_norm, and allow the work to spread across multiple SMs (#15281) |
| 831 | b1afcab804e3281867a5471fbd701e32eb32e512 | 9ef536907de1b50c30e0369284898d30472a755a | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-08-23T15:21:52+02:00 | GitHub | noreply@github.com | 2025-08-23T15:21:52+02:00 | | model : add support for Seed-OSS (#15490) |
| 832 | 9ef536907de1b50c30e0369284898d30472a755a | 21dc4ddaf21b8ed551d717e7606abd2cffbacdbf | Johannes Gäßler | johannesg@5d6.de | 2025-08-23T12:58:58+02:00 | GitHub | noreply@github.com | 2025-08-23T13:58:58+03:00 | | scripts: fix compare-llama-bench.py (#15521) |
| 833 | 21dc4ddaf21b8ed551d717e7606abd2cffbacdbf | 289bf4113ef5c02d8f5eb0cf2d86683d8b8bc4d9 | LaffeyNyaa | 112215776+LaffeyNyaa@users.noreply.github.com | 2025-08-23T16:38:30+08:00 | GitHub | noreply@github.com | 2025-08-23T10:38:30+02:00 | | chat : fix debug build assertion in trim function (#15520) |
| 834 | 289bf4113ef5c02d8f5eb0cf2d86683d8b8bc4d9 | b55f06e1aa67fb10e89f53e31bbccf37eb2678ea | Jeff Bolz | jbolz@nvidia.com | 2025-08-23T02:33:36-05:00 | GitHub | noreply@github.com | 2025-08-23T09:33:36+02:00 | | vulkan: Rewrite synchronization to allow some overlap between nodes (#15489) |
| 835 | b55f06e1aa67fb10e89f53e31bbccf37eb2678ea | 0a9b43e507a359ca392c037cf341f55137ad0b69 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-23T14:58:57+08:00 | GitHub | noreply@github.com | 2025-08-23T08:58:57+02:00 | | vulkan.Dockerfile: install vulkan SDK using tarball (#15282) |
| 836 | 0a9b43e507a359ca392c037cf341f55137ad0b69 | 330c3d2d21b55bca5517db7d2eea2ea8f131df4a | Acly | aclysia@gmail.com | 2025-08-23T08:35:21+02:00 | GitHub | noreply@github.com | 2025-08-23T08:35:21+02:00 | | vulkan : support ggml_mean (#15393) |
| 837 | 330c3d2d21b55bca5517db7d2eea2ea8f131df4a | e92734d51bcb82cc35f0a6b5a14928f0036b2c90 | Jeff Bolz | jbolz@nvidia.com | 2025-08-23T01:31:54-05:00 | GitHub | noreply@github.com | 2025-08-23T08:31:54+02:00 | | vulkan: optimize mul_mat_id loading row ids into shared memory (#15427) |
| 838 | e92734d51bcb82cc35f0a6b5a14928f0036b2c90 | 45363632cbd593537d541e81b600242e0b3d47fc | Johannes Gäßler | johannesg@5d6.de | 2025-08-22T23:47:01+02:00 | GitHub | noreply@github.com | 2025-08-22T23:47:01+02:00 | | test-opt: allow slight inprecision (#15503) |
| 839 | 45363632cbd593537d541e81b600242e0b3d47fc | 32732f2459a598606055f0403f0e4ec148d06d68 | Reese Levine | reeselevine1@gmail.com | 2025-08-22T11:28:03-07:00 | GitHub | noreply@github.com | 2025-08-22T11:28:03-07:00 | | ggml WebGPU: add support for quantization types (#15440) |
| 840 | 32732f2459a598606055f0403f0e4ec148d06d68 | 92f7f0a53cf6484e16a8084ed90807c35a164809 | Aldehir Rojas | hello@alde.dev | 2025-08-22T11:04:08-05:00 | GitHub | noreply@github.com | 2025-08-22T11:04:08-05:00 | | model : gpt-oss add response_format support (#15494) |
| 841 | 92f7f0a53cf6484e16a8084ed90807c35a164809 | b1ab91821f980f8993423c3f2a82a0a0f60c09d2 | rmatif | rmatif@proton.me | 2025-08-22T15:33:15+02:00 | GitHub | noreply@github.com | 2025-08-22T15:33:15+02:00 | | ggml: add `conv3d` op (#15182) |
| 842 | b1ab91821f980f8993423c3f2a82a0a0f60c09d2 | 9ebebef62fd0adf8685874f154e227ea87b7c6f4 | Yavor Ivanov | yavorgenadiev@gmail.com | 2025-08-22T14:06:29+03:00 | GitHub | noreply@github.com | 2025-08-22T13:06:29+02:00 | | cuda : add Pad Reflect 1D support (#14659) |
| 843 | 9ebebef62fd0adf8685874f154e227ea87b7c6f4 | ad5c975c2d0297124fad210776ef8eed6b90d578 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-22T12:22:13+03:00 | GitHub | noreply@github.com | 2025-08-22T12:22:13+03:00 | | llama : remove KV cache defragmentation logic (#15473) |
| 844 | ad5c975c2d0297124fad210776ef8eed6b90d578 | 4afb0a746f22abaa545b3ebdb76a400d7da3a713 | Aaron Teo | aaron.teo1@ibm.com | 2025-08-22T16:11:04+08:00 | GitHub | noreply@github.com | 2025-08-22T16:11:04+08:00 | | ggml-cpu: Support Q5_0 and Q5_1 on s390x (#15486) |
| 845 | 4afb0a746f22abaa545b3ebdb76a400d7da3a713 | e288693669cf9d0a71e2f2b8bd57305f06340257 | 65a | 10104049+65a@users.noreply.github.com | 2025-08-22T08:10:14Z | GitHub | noreply@github.com | 2025-08-22T10:10:14+02:00 | | server : Support multimodal completion and embeddings prompts in JSON format (#15108) |
| 846 | e288693669cf9d0a71e2f2b8bd57305f06340257 | a0f98dd604d34826eb5ea2560d1e23fe726921df | Tarek Dakhran | tarek@liquid.ai | 2025-08-22T09:29:08+02:00 | GitHub | noreply@github.com | 2025-08-22T09:29:08+02:00 | | readme : model : mtdm : lfm2 improvements (#15476) |
| 847 | a0f98dd604d34826eb5ea2560d1e23fe726921df | 54a241f505d515d625767b993bfd573ecee306b9 | Chenguang Li | 757486878@qq.com | 2025-08-22T14:12:07+08:00 | GitHub | noreply@github.com | 2025-08-22T14:12:07+08:00 | | CANN: Optimize RMS_NORM using cache (#15419) |
| 848 | 54a241f505d515d625767b993bfd573ecee306b9 | cd36b5e5c7fed2a3ac671dd542d579ca40b48b54 | Diego Devesa | slarengh@gmail.com | 2025-08-21T14:09:32-07:00 | GitHub | noreply@github.com | 2025-08-21T23:09:32+02:00 | | sched : fix possible use of wrong ids tensor when offloading moe prompt processing (#15488) |
| 849 | cd36b5e5c7fed2a3ac671dd542d579ca40b48b54 | 3f196be84b1376945737163e35a91e29e3e24d2f | Georgi Gerganov | ggerganov@gmail.com | 2025-08-21T19:13:45+03:00 | GitHub | noreply@github.com | 2025-08-21T19:13:45+03:00 | | llama : remove deprecated llama_kv_self API (#15472) |
| 850 | 3f196be84b1376945737163e35a91e29e3e24d2f | 97ae5961a4d9ff9c60f51bac304b97de18e75eaf | Georgi Gerganov | ggerganov@gmail.com | 2025-08-21T18:44:45+03:00 | GitHub | noreply@github.com | 2025-08-21T18:44:45+03:00 | | graph : remove build_attn_with_sinks overload (#15469) |
| 851 | 97ae5961a4d9ff9c60f51bac304b97de18e75eaf | 20c2dac8c6e05f2ad5295ddef1aebaf2d266090e | Acly | aclysia@gmail.com | 2025-08-21T17:01:51+02:00 | GitHub | noreply@github.com | 2025-08-21T17:01:51+02:00 | | vulkan : support conv_2d_dw with f16 weights (#15392) |
| 852 | 20c2dac8c6e05f2ad5295ddef1aebaf2d266090e | 96452a3fa426de83494fb0268a636aa1bfe557fe | Dong Won Kim | 63934649+ddwkim@users.noreply.github.com | 2025-08-22T00:00:16+09:00 | GitHub | noreply@github.com | 2025-08-21T17:00:16+02:00 | | vulkan: add exp operation (#15456) |
| 853 | 96452a3fa426de83494fb0268a636aa1bfe557fe | 9ad5e60dba38a6718366b7ac43e7d8e8abdc36c9 | Jeff Bolz | jbolz@nvidia.com | 2025-08-21T09:55:00-05:00 | GitHub | noreply@github.com | 2025-08-21T16:55:00+02:00 | | vulkan: Reuse conversion results in prealloc_y (#15410) |
| 854 | 9ad5e60dba38a6718366b7ac43e7d8e8abdc36c9 | 715a6db02ccb16284837885f2c6fab05d8f7a6ee | Jie Fu (傅杰) | jiefu@tencent.com | 2025-08-21T22:53:13+08:00 | GitHub | noreply@github.com | 2025-08-21T16:53:13+02:00 | | examples : fix some typos in examples/model-conversion/README.md (#15477) |
| 855 | ad294df03ff2dccd227c3fee653166f3d78b23a4 | 029bb39eb1e7ae0e5df817ce97931abcd5fa52a9 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-08-21T21:42:34+08:00 | GitHub | noreply@github.com | 2025-08-21T15:42:34+02:00 | | examples : install torch-cpu for model conversion tool/example (#15475) |
| 856 | 029bb39eb1e7ae0e5df817ce97931abcd5fa52a9 | 30649cab657d87ac46692332a76e1b75d5d22e00 | Ali Tariq | alitariq4589@gmail.com | 2025-08-21T17:52:16+05:00 | GitHub | noreply@github.com | 2025-08-21T14:52:16+02:00 | | ci : enable RVV1.0 native build (#15386) |
| 857 | 30649cab657d87ac46692332a76e1b75d5d22e00 | 2758fa10dab0556e6c3f130e664750fd6773dc7c | Georgi Gerganov | ggerganov@gmail.com | 2025-08-21T13:42:55+03:00 | GitHub | noreply@github.com | 2025-08-21T13:42:55+03:00 | | ci : continue file download with wget (#15471) |
| 858 | 2758fa10dab0556e6c3f130e664750fd6773dc7c | b108e429043ee5c9fc8fa4957a0a52c3e490d5c9 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-21T12:16:54+02:00 | GitHub | noreply@github.com | 2025-08-21T12:16:54+02:00 | | examples : add model conversion tool/example (#15455) |
| 859 | b108e429043ee5c9fc8fa4957a0a52c3e490d5c9 | 245be739df942861ddc52331b095b40f18e2a3f1 | Michael Giba | michaelgiba@gmail.com | 2025-08-21T05:06:46-05:00 | GitHub | noreply@github.com | 2025-08-21T12:06:46+02:00 | | ci : fix -Werror=return-type in clip.cpp so ci/run.sh can run without issue (#15221) |
| 860 | 245be739df942861ddc52331b095b40f18e2a3f1 | b2caf67db1208fd38a0570785c39f7370d906d8a | Copilot | 198982749+Copilot@users.noreply.github.com | 2025-08-21T11:47:52+02:00 | GitHub | noreply@github.com | 2025-08-21T11:47:52+02:00 | | ci : add copilot-instructions.md (#15286) |
| 861 | b2caf67db1208fd38a0570785c39f7370d906d8a | 2f3dbffb17ef782edfd50e5a130cec6e8a7e47f8 | Julien Denize | 40604584+juliendenize@users.noreply.github.com | 2025-08-21T11:19:50+02:00 | GitHub | noreply@github.com | 2025-08-21T11:19:50+02:00 | | convert : make Mistral community chat templates optional via parameter (#15420) |
| 862 | 2f3dbffb17ef782edfd50e5a130cec6e8a7e47f8 | 945e1f12a6b586ebf82fa4fd7f347225e58174c5 | Jie Fu (傅杰) | jiefu@tencent.com | 2025-08-21T16:54:34+08:00 | GitHub | noreply@github.com | 2025-08-21T11:54:34+03:00 | | common : fix incorrect print of non-ascii characters in the logging (#15466) |
| 863 | 945e1f12a6b586ebf82fa4fd7f347225e58174c5 | 1b0db8f6e08951969e2447c2d18bf638effb8f75 | Xuan-Son Nguyen | son@huggingface.co | 2025-08-21T07:32:26+02:00 | GitHub | noreply@github.com | 2025-08-21T08:32:26+03:00 | | ggml : fix condition of im2col on Metal backend (#15460) |
| 864 | 1b0db8f6e08951969e2447c2d18bf638effb8f75 | 29f538ac630d6544406a0702476e36808a6bd1b3 | stduhpf | stephduh@live.fr | 2025-08-21T07:19:22+02:00 | GitHub | noreply@github.com | 2025-08-21T08:19:22+03:00 | | server : fix webui (#15462) |
| 865 | 29f538ac630d6544406a0702476e36808a6bd1b3 | 8ad038c0fdc719ced9fbf921a02cbae9ad79287f | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-21T06:12:28+02:00 | GitHub | noreply@github.com | 2025-08-21T06:12:28+02:00 | | examples : remove references to `make` in examples [no ci] (#15457) |
| 866 | 8ad038c0fdc719ced9fbf921a02cbae9ad79287f | 5682a3745f2b653dcb855d5766d8edc318fb3336 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-21T11:06:05+08:00 | GitHub | noreply@github.com | 2025-08-21T11:06:05+08:00 | | musa: add GGML_UNUSED_VARS (#15446) |
| 867 | 5682a3745f2b653dcb855d5766d8edc318fb3336 | 1bc664a26a1d93a48baf2483a7ff95291ca8bcb8 | Diego Devesa | slarengh@gmail.com | 2025-08-20T16:35:28-07:00 | GitHub | noreply@github.com | 2025-08-21T01:35:28+02:00 | | sched : copy only the used experts when offloading prompt processing (#15346) |
| 868 | 1bc664a26a1d93a48baf2483a7ff95291ca8bcb8 | 13aeb7aef284daed4ad07dc14b4b3e2a42b5ea97 | teo | TeoZosa@users.noreply.github.com | 2025-08-21T07:10:08+09:00 | GitHub | noreply@github.com | 2025-08-21T00:10:08+02:00 | | server: fix OpenAI API compatibility for usage statistics in chat streams (#15444) |
| 869 | 13aeb7aef284daed4ad07dc14b4b3e2a42b5ea97 | 7a6e91ad26160dd6dfb33d29ac441617422f28e7 | Johannes Gäßler | johannesg@5d6.de | 2025-08-20T23:14:14+02:00 | GitHub | noreply@github.com | 2025-08-20T23:14:14+02:00 | | CUDA: refactor FA support/selection code (#15454) |
| 870 | 7a6e91ad26160dd6dfb33d29ac441617422f28e7 | fec9519802ae1567048abb126cdd5ea160a22d0f | Johannes Gäßler | johannesg@5d6.de | 2025-08-20T16:58:49+02:00 | GitHub | noreply@github.com | 2025-08-20T16:58:49+02:00 | | CUDA: replace GGML_CUDA_F16 with CUDA arch checks (#15433) |
| 871 | fec9519802ae1567048abb126cdd5ea160a22d0f | 657b8a77bd01854f99d37a47318fa24f2e7e298f | Jeff Bolz | jbolz@nvidia.com | 2025-08-20T09:33:14-05:00 | GitHub | noreply@github.com | 2025-08-20T16:33:14+02:00 | | vulkan: shorten pipeline name strings (#15431) |
| 872 | 657b8a77bd01854f99d37a47318fa24f2e7e298f | ec5ab1a36c11dd3efcf4ec8d1ac89a13a8117bc3 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-20T14:26:01+02:00 | GitHub | noreply@github.com | 2025-08-20T14:26:01+02:00 | | chat: handle gpt-oss return/end token inconsistency (#15421) |
| 873 | ec5ab1a36c11dd3efcf4ec8d1ac89a13a8117bc3 | 1a99c2d948209d9ea5eac8b6dc0a297107244540 | Jie Fu (傅杰) | fujie_email@sina.com | 2025-08-20T18:33:30+08:00 | GitHub | noreply@github.com | 2025-08-20T13:33:30+03:00 | | common : fix context shift help message (#15448) |
| 874 | 1a99c2d948209d9ea5eac8b6dc0a297107244540 | 37f10f955f70e0158d50343d0b9a3f92d194daae | xiaobing318 | 71554036+xiaobing318@users.noreply.github.com | 2025-08-20T18:32:05+08:00 | GitHub | noreply@github.com | 2025-08-20T13:32:05+03:00 | | cmake : fix target include directories (#15450) |
| 875 | 37f10f955f70e0158d50343d0b9a3f92d194daae | 2f37014073f4c6ddc8f241c927db87337c71aa52 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-20T12:31:16+02:00 | GitHub | noreply@github.com | 2025-08-20T13:31:16+03:00 | | make : remove make in favor of CMake (#15449) |
| 876 | 2f37014073f4c6ddc8f241c927db87337c71aa52 | a094f381432d92c4bf92d2d6167284316ba73a62 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-20T13:30:46+03:00 | GitHub | noreply@github.com | 2025-08-20T13:30:46+03:00 | | lookahead : add sample command to readme (#15447) |
| 877 | a094f381432d92c4bf92d2d6167284316ba73a62 | fb22dd07a639e81c7415e30b146f545f1a2f2caf | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-20T10:17:37+08:00 | GitHub | noreply@github.com | 2025-08-20T10:17:37+08:00 | | musa: fix build warnings (#15258) |
| 878 | fb22dd07a639e81c7415e30b146f545f1a2f2caf | 9ef6b0b835450ee10cd7be934ad8aef681dc1f43 | lhez | lih@qti.qualcomm.com | 2025-08-20T02:25:51+08:00 | GitHub | noreply@github.com | 2025-08-19T11:25:51-07:00 | | opencl: mark `argsort` unsupported if cols exceed workgroup limit (#15375) |
| 879 | 9ef6b0b835450ee10cd7be934ad8aef681dc1f43 | 1e19f5d462b6df490a13103ca555d851f47e5fa5 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-19T19:58:28+03:00 | GitHub | noreply@github.com | 2025-08-19T19:58:28+03:00 | | model : add gpt-oss type strings (#15424) |
| 880 | 1e19f5d462b6df490a13103ca555d851f47e5fa5 | d2fcd91cf96b46f4485ce46b4e3a32bf0df37715 | Gian-Carlo Pascutto | gcp@sjeng.org | 2025-08-19T18:58:14+02:00 | GitHub | noreply@github.com | 2025-08-19T19:58:14+03:00 | | common : Add top-nsigma sampler to help globally (#15428) |
| 881 | d2fcd91cf96b46f4485ce46b4e3a32bf0df37715 | a6d3cfe7fa6ea1fb0e1ba8243b731db22ddc0b49 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-19T16:46:37+03:00 | GitHub | noreply@github.com | 2025-08-19T16:46:37+03:00 | | server : disable context shift by default (#15416) |
| 882 | a6d3cfe7fa6ea1fb0e1ba8243b731db22ddc0b49 | 67f09a3a27db443f9870aac87e163dba0d08131e | SHUAI YANG | shuaiyang047@163.com | 2025-08-19T21:28:22+08:00 | GitHub | noreply@github.com | 2025-08-19T21:28:22+08:00 | | CANN: optimize rope operator (#15335) |
| 883 | 67f09a3a27db443f9870aac87e163dba0d08131e | 6424594c56f4dbd0573455d89a0d89a0ac093d13 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-19T18:33:47+08:00 | GitHub | noreply@github.com | 2025-08-19T12:33:47+02:00 | | musa: handle __hgt2_mask, available starting from MUSA SDK rc4.3.0 (#15413) |
| 884 | 6424594c56f4dbd0573455d89a0d89a0ac093d13 | e9288e886970884f288533cd597b3798995b4099 | Marvin Gießing | marvin.giessing@gmail.com | 2025-08-19T10:54:31+02:00 | GitHub | noreply@github.com | 2025-08-19T11:54:31+03:00 | | ggml-cpu: add mxfp4 VSX intrinsics for Power9+ (ppc64le) hardware (#15385) |
| 885 | e9288e886970884f288533cd597b3798995b4099 | 9d262f4bad0d37838100133537aaf0a83835ed12 | Xuan-Son Nguyen | son@huggingface.co | 2025-08-19T10:29:36+02:00 | GitHub | noreply@github.com | 2025-08-19T10:29:36+02:00 | | chat : clarify the meaning of reasoning_format (#15408) |
| 886 | 9d262f4bad0d37838100133537aaf0a83835ed12 | f0d3c7405c323784a60f14ddddfbac3f7404d417 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-19T08:45:26+03:00 | GitHub | noreply@github.com | 2025-08-19T08:45:26+03:00 | | server : remove swa_full warning (#15399) |
| 887 | f0d3c7405c323784a60f14ddddfbac3f7404d417 | f08c4c0d8d0cb6caaf8b7ad316039232b9fa059c | Georgi Gerganov | ggerganov@gmail.com | 2025-08-19T08:45:12+03:00 | GitHub | noreply@github.com | 2025-08-19T08:45:12+03:00 | | batched-bench : use rand tokens (#15398) |
| 888 | f08c4c0d8d0cb6caaf8b7ad316039232b9fa059c | 6d7f1117e3e3285d0c5c11b5ebb0439e27920082 | Xuan-Son Nguyen | son@huggingface.co | 2025-08-18T22:53:52+02:00 | GitHub | noreply@github.com | 2025-08-18T22:53:52+02:00 | | mtmd : clean up clip_n_output_tokens (#15391) |
| 889 | 6d7f1117e3e3285d0c5c11b5ebb0439e27920082 | 60212f1ead2dce9bf1ac69633a7069258ae604d8 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T22:02:50+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T22:06:44+03:00 | | codeowners : remove mmv.* |
| 890 | 60212f1ead2dce9bf1ac69633a7069258ae604d8 | f0c541d315e97b297b3421c52ebde53340ee66b3 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T22:02:11+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T22:06:44+03:00 | | sync : ggml |
| 891 | f0c541d315e97b297b3421c52ebde53340ee66b3 | baa9255a45105d2d3b4ec432af13b7a6eda3ff35 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T20:35:47+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T22:06:44+03:00 | | scripts : update sync scripts |
| 892 | baa9255a45105d2d3b4ec432af13b7a6eda3ff35 | 3007baf201e7ffcda17dbdb0335997fa50a6595b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-18T19:30:17+02:00 | GitHub | noreply@github.com | 2025-08-18T19:30:17+02:00 | | llama : merge conts and reshapes and remove unnecessary cont (#15380) |
| 893 | 3007baf201e7ffcda17dbdb0335997fa50a6595b | d1d82416006e7ff41780cb0e9b5f28d30a267497 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-18T18:11:44+03:00 | GitHub | noreply@github.com | 2025-08-18T18:11:44+03:00 | | readme : update hot topics (#15397) |
| 894 | d1d82416006e7ff41780cb0e9b5f28d30a267497 | 618575c5825d7d4f170e686e772178d2aae148ae | davidef | davidef1986@gmail.com | 2025-08-18T16:51:42+02:00 | GitHub | noreply@github.com | 2025-08-18T17:51:42+03:00 | | server : fix incoming tasks not process in order (#15395) |
| 895 | 618575c5825d7d4f170e686e772178d2aae148ae | f44f7931729022c57319a0124931120a169e0da9 | Dobri Danchev | 12420863+danchev@users.noreply.github.com | 2025-08-18T05:50:48-05:00 | GitHub | noreply@github.com | 2025-08-18T12:50:48+02:00 | | Fix broken build: require updated pip to support --break-system-packages (#15357) |
| 896 | f44f7931729022c57319a0124931120a169e0da9 | ae532eac2c1df1d8edc3d2719145895b966de1bf | compilade | git@compilade.net | 2025-08-18T03:23:56-04:00 | GitHub | noreply@github.com | 2025-08-18T09:23:56+02:00 | | ggml-quants : fix make_qp_quants NANs and IQ1 assertion errors (#15379) |
| 897 | ae532eac2c1df1d8edc3d2719145895b966de1bf | e5155e698645242d4f019267ecc40ea9bad81b09 | Jeff Bolz | jbolz@nvidia.com | 2025-08-18T00:56:29-05:00 | GitHub | noreply@github.com | 2025-08-18T07:56:29+02:00 | | vulkan: disable spirv-opt for bfloat16 shaders (#15352) |
| 898 | e5155e698645242d4f019267ecc40ea9bad81b09 | 21c17b5befc5f6be5992bc87fc1ba99d388561df | Oleksandr Kuvshynov | 661042+okuvshynov@users.noreply.github.com | 2025-08-17T18:28:58-04:00 | GitHub | noreply@github.com | 2025-08-18T00:28:58+02:00 | | server : export max observed n_past value (#15361) |
| 899 | 21c17b5befc5f6be5992bc87fc1ba99d388561df | 19f4decae0ead52debe56095ba8d693b4f14e4df | Jeff Bolz | jbolz@nvidia.com | 2025-08-17T11:08:57-05:00 | GitHub | noreply@github.com | 2025-08-17T18:08:57+02:00 | | vulkan: Use larger workgroups for mul_mat_vec when M is small (#15355) |
| 900 | 19f4decae0ead52debe56095ba8d693b4f14e4df | 4d196981d4db79e0105b939eaa7ecd40385b721c | Dong Won Kim | 63934649+ddwkim@users.noreply.github.com | 2025-08-17T23:03:09+09:00 | GitHub | noreply@github.com | 2025-08-17T16:03:09+02:00 | | vulkan: support sqrt (#15370) |
| 901 | 4d196981d4db79e0105b939eaa7ecd40385b721c | b143fbc87af0324aa49e16cd91faf3ba8bb22231 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-17T14:47:42+02:00 | GitHub | noreply@github.com | 2025-08-17T14:47:42+02:00 | | convert : force patch_embd weights to F16 or F32 to avoid broken GGUFs (#15367) |
| 902 | b143fbc87af0324aa49e16cd91faf3ba8bb22231 | de5627910df74298c998e6bb36ee3217375a5719 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-17T13:30:23+02:00 | GitHub | noreply@github.com | 2025-08-17T13:30:23+02:00 | | ci : fix hang in windows-hip build/release (#15365) |
| 903 | de5627910df74298c998e6bb36ee3217375a5719 | 65349f26f2299e06477ec8e85e46243046801358 | Jeff Bolz | jbolz@nvidia.com | 2025-08-17T03:41:45-05:00 | GitHub | noreply@github.com | 2025-08-17T10:41:45+02:00 | | vulkan: Optimize argsort (#15354) |
| 904 | 65349f26f2299e06477ec8e85e46243046801358 | 1fe00296f587dfca0957e006d146f5875b61e43d | Tarek Dakhran | t.dakhran@gmail.com | 2025-08-16T23:33:54+02:00 | GitHub | noreply@github.com | 2025-08-16T23:33:54+02:00 | | model : support vision LiquidAI LFM2-VL family (#15347) |
| 905 | 1fe00296f587dfca0957e006d146f5875b61e43d | de2192794f4e8e04f2e8167ef2424905145e88fc | Jeff Bolz | jbolz@nvidia.com | 2025-08-16T11:48:22-05:00 | GitHub | noreply@github.com | 2025-08-16T11:48:22-05:00 | | vulkan: fuse adds (#15252) |
| 906 | de2192794f4e8e04f2e8167ef2424905145e88fc | 2e2b22ba6607414a5d619ac6d2f034b5b02214e5 | Jeff Bolz | jbolz@nvidia.com | 2025-08-16T04:18:31-05:00 | GitHub | noreply@github.com | 2025-08-16T11:18:31+02:00 | | vulkan: Support mul_mat_id with f32 accumulators (#15337) |
| 907 | 2e2b22ba6607414a5d619ac6d2f034b5b02214e5 | 912ff8c119f01ae029543c7fdf7a84f91a0437a3 | Jeff Bolz | jbolz@nvidia.com | 2025-08-16T03:58:38-05:00 | GitHub | noreply@github.com | 2025-08-16T10:58:38+02:00 | | vulkan: Add missing bounds checking to scalar/coopmat1 mul_mat_id (#15334) |
| 908 | 912ff8c119f01ae029543c7fdf7a84f91a0437a3 | 5e6229a8409ac786e62cb133d09f1679a9aec13e | rmatif | kingrealriadh@gmail.com | 2025-08-16T10:05:55+02:00 | GitHub | noreply@github.com | 2025-08-16T01:05:55-07:00 | | OpenCL: add initial FA support (#14987) |
| 909 | 5e6229a8409ac786e62cb133d09f1679a9aec13e | e2c1bfff5305c661ac53e9d57cb732ff626a2242 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-15T19:50:52+02:00 | GitHub | noreply@github.com | 2025-08-15T19:50:52+02:00 | | common : fix double bos, use common_chat_templates for add_bos and add_eos (#15326) |
| 910 | e2c1bfff5305c661ac53e9d57cb732ff626a2242 | 5edf1592fdb9131d01321aeef4241c6a34969e27 | lhez | lih@qti.qualcomm.com | 2025-08-16T00:52:14+08:00 | GitHub | noreply@github.com | 2025-08-15T09:52:14-07:00 | | opencl: add initial mxfp4 support via mv (#15270) |
| 911 | 5edf1592fdb9131d01321aeef4241c6a34969e27 | db3010bd23980bad5f5d93ec3e6757ec531a413b | Georgi Gerganov | ggerganov@gmail.com | 2025-08-15T17:16:36+03:00 | GitHub | noreply@github.com | 2025-08-15T16:16:36+02:00 | | vulkan : fix out-of-bounds access in argmax kernel (#15342) |
| 912 | db3010bd23980bad5f5d93ec3e6757ec531a413b | ff27f80a74bbe5303acd511a6781a1de6d619b3c | Georgi Gerganov | ggerganov@gmail.com | 2025-08-15T16:28:28+03:00 | GitHub | noreply@github.com | 2025-08-15T15:28:28+02:00 | | vulkan : fix compile warnings on macos (#15340) |
| 913 | ff27f80a74bbe5303acd511a6781a1de6d619b3c | d3248d9b6557c75d59954c594bb53cf517591e91 | Aaron Teo | aaron.teo1@ibm.com | 2025-08-15T21:11:22+08:00 | GitHub | noreply@github.com | 2025-08-15T21:11:22+08:00 | | ggml: initial IBM zDNN backend (#14975) |
| 914 | d3248d9b6557c75d59954c594bb53cf517591e91 | 7aeee88cfe9c0434eac722b0f7c21404f48758a5 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-15T14:02:39+02:00 | GitHub | noreply@github.com | 2025-08-15T14:02:39+02:00 | | ci : fix ios-xcode-build (#15324) |
| 915 | 7aeee88cfe9c0434eac722b0f7c21404f48758a5 | b07791aa1d4831a08ad54ca19aa206c0c5c0f34a | Diego Devesa | slarengh@gmail.com | 2025-08-15T03:27:02-07:00 | GitHub | noreply@github.com | 2025-08-15T12:27:02+02:00 | | ci : move ccache action to ggml-org fork (#15328) |
| 916 | b07791aa1d4831a08ad54ca19aa206c0c5c0f34a | 4227c9be4268ac844921b90f31595f81236bd317 | Johannes Gäßler | johannesg@5d6.de | 2025-08-15T11:23:17+02:00 | GitHub | noreply@github.com | 2025-08-15T11:23:17+02:00 | | test-opt: fix backend support check (#15317) |
| 917 | 4227c9be4268ac844921b90f31595f81236bd317 | df36bce667bf14f8e538645547754386f9516326 | Johannes Gäßler | johannesg@5d6.de | 2025-08-14T23:21:24+02:00 | GitHub | noreply@github.com | 2025-08-14T23:21:24+02:00 | | CUDA: fix negative KV_max values in FA (#15321) |
| 918 | df36bce667bf14f8e538645547754386f9516326 | f75b8306472811ecc4e4457d5573a09201f02183 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T22:10:51+03:00 | GitHub | noreply@github.com | 2025-08-14T22:10:51+03:00 | | eval-callback : stop on first NaN (#15320) |
| 919 | f75b8306472811ecc4e4457d5573a09201f02183 | 7a0de960452f9a57de7f1d167e57b6f3ee5ac1b6 | Diego Devesa | slarengh@gmail.com | 2025-08-14T10:28:29-07:00 | GitHub | noreply@github.com | 2025-08-14T10:28:29-07:00 | | chat : include kwargs in template example (#15309) |
| 920 | 7a0de960452f9a57de7f1d167e57b6f3ee5ac1b6 | e4e915912cfd2ee15c5a4a0074813232134892f6 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-14T17:56:26+02:00 | GitHub | noreply@github.com | 2025-08-14T17:56:26+02:00 | | llama : add 18-layer model type for Gemma 3-270m (#15319) |
| 921 | e4e915912cfd2ee15c5a4a0074813232134892f6 | 5ba36f61033d2819be27e2abb70e9a3bd20c0fde | simevo | github@simevo.com | 2025-08-14T17:45:27+02:00 | GitHub | noreply@github.com | 2025-08-14T18:45:27+03:00 | | devops : fix compile bug when the BASE_CUDA_DEV_CONTAINER is based on Ubuntu 24.04 (#15005) |
| 922 | 5ba36f61033d2819be27e2abb70e9a3bd20c0fde | b204a5a234f4cf0b3f449476da639a5510c6d157 | uvos | carl@uvos.xyz | 2025-08-14T16:23:56+02:00 | GitHub | noreply@github.com | 2025-08-14T16:23:56+02:00 | | HIP: Cleanup hipification header (#15285) |
| 923 | b204a5a234f4cf0b3f449476da639a5510c6d157 | 646944cfa8961afd914dd6637739b3cda9a72e11 | Aldehir Rojas | hello@alde.dev | 2025-08-14T09:23:11-05:00 | GitHub | noreply@github.com | 2025-08-14T17:23:11+03:00 | | gpt-oss: implement harmony parsing (#15181) |
| 924 | 646944cfa8961afd914dd6637739b3cda9a72e11 | 1a01899b612ae8f99a174ad076207090e08d4d7b | Christian Kastner | ckk@kvr.at | 2025-08-14T16:22:58+02:00 | GitHub | noreply@github.com | 2025-08-14T16:22:58+02:00 | | docker : Enable GGML_CPU_ALL_VARIANTS for ARM (#15267) |
| 925 | 1a01899b612ae8f99a174ad076207090e08d4d7b | 863d341eeb81db104902b14b6f9413daa515e957 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T17:16:03+03:00 | GitHub | noreply@github.com | 2025-08-14T17:16:03+03:00 | | readme : update hot topics (#15315) |
| 926 | 863d341eeb81db104902b14b6f9413daa515e957 | d32e03f4495d3efa1c5126f53b449f1d429c5664 | Jeff Bolz | jbolz@nvidia.com | 2025-08-14T08:38:10-05:00 | GitHub | noreply@github.com | 2025-08-14T08:38:10-05:00 | | vulkan: perf_logger improvements (#15246) |
| 927 | d32e03f4495d3efa1c5126f53b449f1d429c5664 | 3973163bff40f7f5161b0f08a0011729e2b0406a | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T14:59:50+03:00 | GitHub | noreply@github.com | 2025-08-14T14:59:50+03:00 | | server : add SWA checkpoints (#15293) |
| 928 | 3973163bff40f7f5161b0f08a0011729e2b0406a | 5ade3000bd6040bf6878eafd6c77168552f73c47 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T14:19:23+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T14:59:27+03:00 | | sync : ggml |
| 929 | 5ade3000bd6040bf6878eafd6c77168552f73c47 | 8b2483730f90d6f23e2b188a54a210498bba937f | Jason Ni | jason.ni.py@gmail.com | 2025-08-14T19:17:51+08:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T14:59:27+03:00 | | ggml: fix ggml_conv_1d_dw bug (ggml/1323) |
| 930 | 8b2483730f90d6f23e2b188a54a210498bba937f | 810b9fc8b99dd55517a3e94c6dee29748d482559 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T13:41:03+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-14T14:59:27+03:00 | | tests : remove unused includes (ggml/0) |
| 931 | 810b9fc8b99dd55517a3e94c6dee29748d482559 | 4ebd0c125b24a0d7a78b0ffc1d9567530ed8f0c4 | kallewoof | karljohan-alm@garage.co.jp | 2025-08-14T20:03:30+09:00 | GitHub | noreply@github.com | 2025-08-14T14:03:30+03:00 | | perplexity : provide a helpful hint for has_cpl case in split_equal error. (#15304) |
| 932 | 4ebd0c125b24a0d7a78b0ffc1d9567530ed8f0c4 | 5cdb27e0917479d2d742cea7beee089574bb09fa | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-14T12:22:07+02:00 | GitHub | noreply@github.com | 2025-08-14T13:22:07+03:00 | | cuda : fix GGML_CUDA_GRAPHS=OFF (#15300) |
| 933 | 5cdb27e0917479d2d742cea7beee089574bb09fa | 3ea913f1ce9567289aedd866a569dbab8fb8e419 | Jonathan Graehl | 99024+graehl@users.noreply.github.com | 2025-08-14T03:03:57-07:00 | GitHub | noreply@github.com | 2025-08-14T12:03:57+02:00 | | finetune: SGD optimizer, more CLI args (#13873) |
| 934 | 3ea913f1ce9567289aedd866a569dbab8fb8e419 | 29c8fbe4e05fd23c44950d0958299e25fbeabc5c | kallewoof | karljohan-alm@garage.co.jp | 2025-08-14T15:16:32+09:00 | GitHub | noreply@github.com | 2025-08-14T09:16:32+03:00 | | perplexity: give more information about constraints on failure (#15303) |
| 935 | 29c8fbe4e05fd23c44950d0958299e25fbeabc5c | 1adc9812bd33dc85489bf093528d61c22917d54f | uvos | carl@uvos.xyz | 2025-08-13T20:44:30+02:00 | GitHub | noreply@github.com | 2025-08-13T20:44:30+02:00 | | HIP: bump requirement to rocm 6.1 (#15296) |
| 936 | 1adc9812bd33dc85489bf093528d61c22917d54f | b3e16665e14032d1276f9d0263acef8321b6f518 | Bas Nijholt | basnijholt@gmail.com | 2025-08-13T11:21:31-07:00 | GitHub | noreply@github.com | 2025-08-13T11:21:31-07:00 | | fix(nix): remove non-functional llama-cpp cachix cache from flake.nix (#15295) |
| 937 | b3e16665e14032d1276f9d0263acef8321b6f518 | c24f4e26883a91219af5acfaf1471ca7ef582683 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-13T15:43:00+02:00 | GitHub | noreply@github.com | 2025-08-13T15:43:00+02:00 | | server : enable -td and -tbd parameters (#15172) |
| 938 | c24f4e26883a91219af5acfaf1471ca7ef582683 | d8914fc47e8a69b28c670325cb1c8ce33e3a2960 | Judd | 4046440+foldl@users.noreply.github.com | 2025-08-13T18:45:15+08:00 | GitHub | noreply@github.com | 2025-08-13T13:45:15+03:00 | | ggml : update `ggml_rope_multi` (#12665) |
| 939 | d8914fc47e8a69b28c670325cb1c8ce33e3a2960 | e885445bc1719d2e289a9b2e11714e0f994e68be | Copilot | 198982749+Copilot@users.noreply.github.com | 2025-08-13T12:44:40+02:00 | GitHub | noreply@github.com | 2025-08-13T12:44:40+02:00 | | common : add --override-tensor-draft, --cpu-moe-draft and --n-cpu-moe-draft parameters (#15191) |
| 940 | e885445bc1719d2e289a9b2e11714e0f994e68be | 648ebcdb739abfb5a9be14f4c5c0fd19ff268ef0 | Aldehir Rojas | hello@alde.dev | 2025-08-13T05:28:21-05:00 | GitHub | noreply@github.com | 2025-08-13T12:28:21+02:00 | | server : filter out harmony thought messages (#15278) |
| 941 | 648ebcdb739abfb5a9be14f4c5c0fd19ff268ef0 | 07aa869a91837d95fcb5612c65a188763ac38647 | Ali Tariq | alitariq4589@gmail.com | 2025-08-13T15:14:44+05:00 | GitHub | noreply@github.com | 2025-08-13T13:14:44+03:00 | | ci : Added CI with RISC-V RVV1.0 Hardware (#14439) |
| 942 | 07aa869a91837d95fcb5612c65a188763ac38647 | 00f35d509e5367bf19ccaf4b4adddc38db4811c6 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-13T11:30:45+02:00 | GitHub | noreply@github.com | 2025-08-13T11:30:45+02:00 | | ci : add more python requirements to copilot-setup-steps (#15289) |
| 943 | 00f35d509e5367bf19ccaf4b4adddc38db4811c6 | 6028bf74351d35a06bd98498624f8c2f029f7d1a | Georgi Gerganov | ggerganov@gmail.com | 2025-08-13T11:09:39+03:00 | GitHub | noreply@github.com | 2025-08-13T11:09:39+03:00 | | ggml : repack block_iq4_nlx8 (#14904) |
| 944 | 6028bf74351d35a06bd98498624f8c2f029f7d1a | bc5182272c373267352bc689e5fca276934bea2d | Oliver Simons | osimons@nvidia.com | 2025-08-13T10:04:46+02:00 | GitHub | noreply@github.com | 2025-08-13T10:04:46+02:00 | | CUDA: Optimize `reduce_rows_f32` kernel, leading up to 25x perf improvement on kernel-level and 10% perf increase for Gemma3n (#15132) |
| 945 | bc5182272c373267352bc689e5fca276934bea2d | e71d48e3265027351e44a8e198f933c98f242c2e | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-13T09:07:13+02:00 | GitHub | noreply@github.com | 2025-08-13T09:07:13+02:00 | | ci : add copilot-setup-steps.yml (#15214) |
| 946 | e71d48e3265027351e44a8e198f933c98f242c2e | b0493156fa8622694f21f42460a84da3eded0bc0 | Tak-RS | snosk.t@gmail.com | 2025-08-13T14:54:30+09:00 | GitHub | noreply@github.com | 2025-08-13T08:54:30+03:00 | | ggml-rpc: chunk send()/recv() to avoid EINVAL for very large tensors over RPC (macOS & others) (#15188) |
| 947 | 3a617cb18121a85eed17ac5585788810c5d9c1e0 | c42293848dd9f03eba9a6c1b85a600b398f06541 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-13T10:58:59+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-13T10:58:59+08:00 | | add calrt_decode |
| 948 | b0493156fa8622694f21f42460a84da3eded0bc0 | f4586ee5986d6f965becb37876d6f3666478a961 | uvos | carl@uvos.xyz | 2025-08-12T22:15:12+02:00 | GitHub | noreply@github.com | 2025-08-12T22:15:12+02:00 | | HIP: disable sync warp shuffel operators from clr amd_warp_sync_functions.h (#15273) |
| 949 | f4586ee5986d6f965becb37876d6f3666478a961 | 60a76588106d22ccacd6b1a0d15d8545861d0a0d | Romain Biessy | romain.biessy@codeplay.com | 2025-08-12T13:58:22+02:00 | GitHub | noreply@github.com | 2025-08-12T13:58:22+02:00 | | sycl: Fix and disable more configurations of mul_mat (#15151) |
| 950 | 60a76588106d22ccacd6b1a0d15d8545861d0a0d | efe3a90996ca6e67ca90f337fc03ad551082c408 | rmatif | kingrealriadh@gmail.com | 2025-08-12T11:42:41+02:00 | GitHub | noreply@github.com | 2025-08-12T02:42:41-07:00 | | opencl: allow mixed f16/f32 `add` (#15140) |
| 951 | efe3a90996ca6e67ca90f337fc03ad551082c408 | bbd57b7eafb3b32e3f7f3a2175bce0b35abc7de8 | Aman Gupta | amangupta052@gmail.com | 2025-08-12T17:21:45+08:00 | GitHub | noreply@github.com | 2025-08-12T17:21:45+08:00 | | CUDA cmake: add `-lineinfo` for easier debug (#15260) |
| 952 | bbd57b7eafb3b32e3f7f3a2175bce0b35abc7de8 | 25ff6f7659f6a5c47d6a73eada5813f0495331f0 | Chenguang Li | 757486878@qq.com | 2025-08-12T16:12:13+08:00 | GitHub | noreply@github.com | 2025-08-12T16:12:13+08:00 | | CANN: GGML_OP_CPY optimization (#15070) |
| 953 | 25ff6f7659f6a5c47d6a73eada5813f0495331f0 | be48528b068111304e4a0bb82c028558b5705f05 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-12T10:02:51+08:00 | GitHub | noreply@github.com | 2025-08-12T10:02:51+08:00 | | musa: fix failures in test-backend-ops for mul_mat_id op (#15236) |
| 954 | be48528b068111304e4a0bb82c028558b5705f05 | cf9e5648a7ae02f7728fefcbcfd2979d83c96e8f | hipudding | huafengchun@gmail.com | 2025-08-11T22:50:31+08:00 | GitHub | noreply@github.com | 2025-08-11T22:50:31+08:00 | | CANN: Add broadcast for softmax and FA (#15208) |
| 955 | cf9e5648a7ae02f7728fefcbcfd2979d83c96e8f | fba5c0d680b555cbac563fef57b33bc5c9621e08 | rainred | 107027757+gryffindor-rr@users.noreply.github.com | 2025-08-11T22:12:12+08:00 | GitHub | noreply@github.com | 2025-08-11T16:12:12+02:00 | | mtmd : Fix MinicpmV model converter and clip to avoid using hardcode. (#14750) |
| 956 | fba5c0d680b555cbac563fef57b33bc5c9621e08 | 53d0a1265826f40c5dbc01d06aeab9e14fcbd69b | Xuan-Son Nguyen | son@huggingface.co | 2025-08-11T15:31:35+02:00 | GitHub | noreply@github.com | 2025-08-11T15:31:35+02:00 | | chat : hotfix gpt-oss jinja raising an exception (#15243) |
| 957 | 53d0a1265826f40c5dbc01d06aeab9e14fcbd69b | 27093afe78912494073eb043fec93a007e49653c | Xuan-Son Nguyen | son@huggingface.co | 2025-08-11T14:48:41+02:00 | GitHub | noreply@github.com | 2025-08-11T14:48:41+02:00 | | server : allow specifying reasoning_format in HTTP request (#15238) |
| 958 | 27093afe78912494073eb043fec93a007e49653c | 228f724d9ce6c56e8cec75bfdabab4dd013def7f | Zagaj | m.zagajewska@gmail.com | 2025-08-11T14:27:54+02:00 | GitHub | noreply@github.com | 2025-08-11T15:27:54+03:00 | | readme : update infra list (#15234) |
| 959 | 228f724d9ce6c56e8cec75bfdabab4dd013def7f | cd3069dfcbeee8e0e96cffb93b0ebb9e595e273a | Georgi Gerganov | ggerganov@gmail.com | 2025-08-11T13:58:24+03:00 | GitHub | noreply@github.com | 2025-08-11T13:58:24+03:00 | | kv-cache : fix seq_rm with seq_id == -1 (#15226) |
| 960 | cd3069dfcbeee8e0e96cffb93b0ebb9e595e273a | 50e81bdf5db563ab57a9a722b08f96fa8a76c927 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-11T11:21:19+02:00 | GitHub | noreply@github.com | 2025-08-11T11:21:19+02:00 | | kv-cache : log (debug) all streams in find_slot (#15176) |
| 961 | 50e81bdf5db563ab57a9a722b08f96fa8a76c927 | 1ebbaddff2a44b0599df659175d3274bd5bbeb81 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-11T11:15:44+02:00 | GitHub | noreply@github.com | 2025-08-11T11:15:44+02:00 | | convert : fix merge conflicts (#15229) |
| 962 | 1ebbaddff2a44b0599df659175d3274bd5bbeb81 | a3a7874272e5a060079658eb5cca4617b7f99062 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-11T10:21:24+02:00 | GitHub | noreply@github.com | 2025-08-11T11:21:24+03:00 | | perplexity : update comments/error msg to use decode [no ci] (#15227) |
| 963 | a3a7874272e5a060079658eb5cca4617b7f99062 | 002cb1bb3345967fbe0fa7766c2d94c2da31ef45 | Julien Denize | 40604584+juliendenize@users.noreply.github.com | 2025-08-11T10:07:49+02:00 | GitHub | noreply@github.com | 2025-08-11T10:07:49+02:00 | | convert : improve Mistral models integration (#14737) |
| 964 | 002cb1bb3345967fbe0fa7766c2d94c2da31ef45 | 79c1160b073b8148a404f3dd2584be1606dccc66 | Charles Xu | charles.xu@arm.com | 2025-08-11T09:59:26+02:00 | GitHub | noreply@github.com | 2025-08-11T09:59:26+02:00 | | kleidiai: fix unsigned overflow bug (#15150) |
| 965 | 79c1160b073b8148a404f3dd2584be1606dccc66 | 34c9d765bf173c551398f1e7fa4595019bc53bab | David Zhao | 90013954+Your-Cheese@users.noreply.github.com | 2025-08-09T13:29:43-05:00 | GitHub | noreply@github.com | 2025-08-09T20:29:43+02:00 | | cuda: refactored ssm_scan and use CUB (#13291) |
| 966 | 34c9d765bf173c551398f1e7fa4595019bc53bab | e54d41befcc1575f4c898c5ff4ef43970cead75f | Aman Gupta | amangupta052@gmail.com | 2025-08-09T20:00:24+08:00 | GitHub | noreply@github.com | 2025-08-09T20:00:24+08:00 | | CUDA: add attention sinks for tile and wmma (#15178) |
| 967 | e54d41befcc1575f4c898c5ff4ef43970cead75f | 4850b52aedceeb70bb4fe49f2d7cd1df6ee98682 | compilade | git@compilade.net | 2025-08-08T17:48:26-04:00 | GitHub | noreply@github.com | 2025-08-08T17:48:26-04:00 | | gguf-py : add Numpy MXFP4 de/quantization support (#15111) |
| 968 | 4850b52aedceeb70bb4fe49f2d7cd1df6ee98682 | cd6983d56d2cce94ecb86bb114ae8379a609073c | Johannes Gäßler | johannesg@5d6.de | 2025-08-08T23:04:36+02:00 | GitHub | noreply@github.com | 2025-08-08T23:04:36+02:00 | | server-bench: external OAI servers, sqlite (#15179) |
| 969 | cd6983d56d2cce94ecb86bb114ae8379a609073c | 6c7e9a54406dbba5e53754a8f70a285414717b06 | AN Long | aisk@users.noreply.github.com | 2025-08-08T21:37:22+09:00 | GitHub | noreply@github.com | 2025-08-08T14:37:22+02:00 | | ggml : fix field name when new ggml_backend (#14944) |
| 970 | 6c7e9a54406dbba5e53754a8f70a285414717b06 | 1425f587a82bc303469b5c32759a2746ba4e1e20 | Olivier Chafik | olivier.chafik@gmail.com | 2025-08-08T10:45:18+01:00 | GitHub | noreply@github.com | 2025-08-08T10:45:18+01:00 | | vendor: sync minja (#15161) |
| 971 | 1425f587a82bc303469b5c32759a2746ba4e1e20 | aaa3d07ae749b781d6135eaff23c7fa8a4ab404a | Johannes Gäßler | johannesg@5d6.de | 2025-08-08T08:19:58+02:00 | GitHub | noreply@github.com | 2025-08-08T08:19:58+02:00 | | CUDA: attention sinks for mma FlashAttention (#15157) |
| 972 | aaa3d07ae749b781d6135eaff23c7fa8a4ab404a | 50aa9389014bba2dd12234132aa6b8ca3601a17f | lhez | lih@qti.qualcomm.com | 2025-08-08T13:47:03+09:00 | GitHub | noreply@github.com | 2025-08-07T21:47:03-07:00 | | opencl: support sink in `soft_max` (attn sinks) (#15152) |
| 973 | 50aa9389014bba2dd12234132aa6b8ca3601a17f | c4f53563df4575196ea13f5ed669ea8ea659a6be | Xuan-Son Nguyen | son@huggingface.co | 2025-08-07T23:26:03+02:00 | GitHub | noreply@github.com | 2025-08-07T23:26:03+02:00 | | convert : support non-mxfp4 HF model (#15153) |
| 974 | c4f53563df4575196ea13f5ed669ea8ea659a6be | a0552c8beef74e843bb085c8ef0c63f9ed7a2b27 | Jeff Bolz | jbolz@nvidia.com | 2025-08-07T15:44:20-05:00 | GitHub | noreply@github.com | 2025-08-07T22:44:20+02:00 | | vulkan: support fattn sinks (#15126) |
| 975 | a0552c8beef74e843bb085c8ef0c63f9ed7a2b27 | 99acbc9921b119aa7ed929eb5780a66a8f06e6d9 | Jeff Bolz | jbolz@nvidia.com | 2025-08-07T15:07:11-05:00 | GitHub | noreply@github.com | 2025-08-07T22:07:11+02:00 | | vulkan: Add env var to disable host visible vidmem (#15109) |
| 976 | 99acbc9921b119aa7ed929eb5780a66a8f06e6d9 | 7ad67ba9fe2b909e271dd31b99c5fce3aba35899 | RunningLeon | mnsheng@yeah.net | 2025-08-08T00:20:40+08:00 | GitHub | noreply@github.com | 2025-08-07T18:20:40+02:00 | | llama : Support intern-s1 (#14875) |
| 977 | 7ad67ba9fe2b909e271dd31b99c5fce3aba35899 | 9a96389544a08fd829fccda28142ce2066017fde | uvos | carl@uvos.xyz | 2025-08-07T16:44:14+02:00 | GitHub | noreply@github.com | 2025-08-07T16:44:14+02:00 | | HIP: add cmake option to enable compiler output of kernel resource usage metrics (#15103) |
| 978 | 9a96389544a08fd829fccda28142ce2066017fde | 1d72c841888b9450916bdd5a9b3274da380f5b36 | Christian Kastner | ckk@kvr.at | 2025-08-07T13:45:41+02:00 | GitHub | noreply@github.com | 2025-08-07T13:45:41+02:00 | | ggml: Skip backend library linking code when GGML_BACKEND_DL=ON (#15094) |
| 979 | 1d72c841888b9450916bdd5a9b3274da380f5b36 | 20638e4f16fcc21f836c7556e83bbf532bb5a0f0 | Johannes Gäßler | johannesg@5d6.de | 2025-08-07T10:53:21+02:00 | GitHub | noreply@github.com | 2025-08-07T10:53:21+02:00 | | CUDA: GEMM for FP32/FP16/BF16 and ne11 <= 16 (#15131) |
| 980 | 20638e4f16fcc21f836c7556e83bbf532bb5a0f0 | 36d3f00e142696f708ab297b9f8f1c825594712d | Johannes Gäßler | johannesg@5d6.de | 2025-08-07T08:50:30+02:00 | GitHub | noreply@github.com | 2025-08-07T08:50:30+02:00 | | scripts: fix crash when --tool is not set (#15133) |
| 981 | 36d3f00e142696f708ab297b9f8f1c825594712d | 5fd160bbd9d70b94b5b11b0001fd7f477005e4a0 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-07T05:31:48+02:00 | GitHub | noreply@github.com | 2025-08-07T05:31:48+02:00 | | requirements : fix PyTorch uint64 compatibility (#15134) |
| 982 | 5fd160bbd9d70b94b5b11b0001fd7f477005e4a0 | 756cfea82608911bbfcbf45164b8fdaddbafaa31 | Reese Levine | reeselevine1@gmail.com | 2025-08-06T15:14:40-07:00 | GitHub | noreply@github.com | 2025-08-06T15:14:40-07:00 | | ggml: Add basic SET_ROWS support in WebGPU (#15137) |
| 983 | 756cfea82608911bbfcbf45164b8fdaddbafaa31 | e725a1a982ca870404a9c4935df52466327bbd02 | rmatif | kingrealriadh@gmail.com | 2025-08-06T23:17:51+02:00 | GitHub | noreply@github.com | 2025-08-06T14:17:51-07:00 | | fix profiling crash (#15072) |
| 984 | e725a1a982ca870404a9c4935df52466327bbd02 | 3db4da56a5a00c4b47bbb3b18f85fb4473662adb | lhez | lih@qti.qualcomm.com | 2025-08-07T04:12:17+09:00 | GitHub | noreply@github.com | 2025-08-06T12:12:17-07:00 | | opencl: add `swiglu_oai` and `add_id` (#15121) |
| 985 | 3db4da56a5a00c4b47bbb3b18f85fb4473662adb | 476aa3fd5779b32d06cd84338f777da82341195c | Sachin Desai | smdesai@gmail.com | 2025-08-06T11:27:30-07:00 | GitHub | noreply@github.com | 2025-08-06T20:27:30+02:00 | | chat : support Granite model reasoning and tool call (#14864) |
| 986 | 476aa3fd5779b32d06cd84338f777da82341195c | 0d8831543cdc368fb248bae6f1b4aa5516684edc | Juk Armstrong | 69222624+jukofyork@users.noreply.github.com | 2025-08-06T17:28:48+01:00 | GitHub | noreply@github.com | 2025-08-06T17:28:48+01:00 | | Fixed name `-override-tensors` to `-override-tensor` (#15129) |
| 987 | 0d8831543cdc368fb248bae6f1b4aa5516684edc | 65c797c4fad4d9966695ac4b4a1560be44109267 | Diego Devesa | slarengh@gmail.com | 2025-08-06T05:37:35-07:00 | GitHub | noreply@github.com | 2025-08-06T14:37:35+02:00 | | ggml : fix fallback to CPU for ununsupported ops (#15118) |
| 988 | 65c797c4fad4d9966695ac4b4a1560be44109267 | 25726898e855ec6dffba227f2233a63c57184036 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-06T13:26:49+02:00 | GitHub | noreply@github.com | 2025-08-06T13:26:49+02:00 | | chat : fix yandex chat template (#15116) |
| 989 | 25726898e855ec6dffba227f2233a63c57184036 | 2241453252147bb7362a286977ee9f9a92130062 | stevenkuang | stevenkuang@tencent.com | 2025-08-06T17:48:30+08:00 | GitHub | noreply@github.com | 2025-08-06T11:48:30+02:00 | | chat : fix hunyuan auto-detection (#15114) |
| 990 | c42293848dd9f03eba9a6c1b85a600b398f06541 | 1cbf442f50a154aac0bf35d7b8b3436a2cdfd6f5 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-06T15:46:06+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-06T15:46:06+08:00 | | sync with calrt API |
| 991 | 1cbf442f50a154aac0bf35d7b8b3436a2cdfd6f5 | ba67d6c03fed3193987a4ef13050e6e2a4451df4 | Yunzhe Jia | yunzhe@calculet.tech | 2025-07-24T15:50:19+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-08-06T15:45:10+08:00 | | 1.sync with calrt prefill,decode,kv_update op 2.add hw pattern untile function |
| 992 | 2241453252147bb7362a286977ee9f9a92130062 | 9515c6131aecaccc955fdedcfe16c3e030aaefcb | Chenguang Li | 757486878@qq.com | 2025-08-06T14:12:42+08:00 | GitHub | noreply@github.com | 2025-08-06T14:12:42+08:00 | | CANN: add support for ACL Graph (#15065) |
| 993 | 9515c6131aecaccc955fdedcfe16c3e030aaefcb | fd1234cb468935ea087d6929b2487926c3afff4b | Reese Levine | reeselevine1@gmail.com | 2025-08-05T16:26:38-07:00 | GitHub | noreply@github.com | 2025-08-05T16:26:38-07:00 | | ggml: WebGPU disable SET_ROWS for now (#15078) |
| 994 | fd1234cb468935ea087d6929b2487926c3afff4b | f324a3b715d5c1081c110ce459f8a8486fb1ee89 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-05T22:10:36+03:00 | GitHub | noreply@github.com | 2025-08-05T22:10:36+03:00 | | llama : add gpt-oss (#15091) |
| 995 | f324a3b715d5c1081c110ce459f8a8486fb1ee89 | be426425817bc3e6a2d91dae476dba6fa85894be | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-05T20:43:36+02:00 | GitHub | noreply@github.com | 2025-08-05T20:43:36+02:00 | | chat : only remove double bos/eos if added (#15086) |
| 996 | be426425817bc3e6a2d91dae476dba6fa85894be | 3306ceabf02e3df66666e5851800e843c7ca207e | Georgi Gerganov | ggerganov@gmail.com | 2025-08-05T20:19:33+03:00 | GitHub | noreply@github.com | 2025-08-05T20:19:33+03:00 | | readme : update hot topics (#15097) |
| 997 | 3306ceabf02e3df66666e5851800e843c7ca207e | c81de6e107ef51ef76aadcb8a6f008711c462517 | Romain Biessy | romain.biessy@codeplay.com | 2025-08-05T18:39:55+02:00 | GitHub | noreply@github.com | 2025-08-05T18:39:55+02:00 | | sycl: fix mul_mat selection (#15092) |
| 998 | c81de6e107ef51ef76aadcb8a6f008711c462517 | 22f060c9c4b5ef49a83a20eda25fcb792419580b | Juk Armstrong | 69222624+jukofyork@users.noreply.github.com | 2025-08-05T13:56:44+01:00 | GitHub | noreply@github.com | 2025-08-05T13:56:44+01:00 | | Fix `glm4moe` bug (#15088) |
| 999 | 22f060c9c4b5ef49a83a20eda25fcb792419580b | ee3a9fcf88fe5b5e1213711e05861b83cd4fdfe6 | Alex Wu | dindinw@users.noreply.github.com | 2025-08-05T19:56:44+08:00 | GitHub | noreply@github.com | 2025-08-05T13:56:44+02:00 | | webui: fix markdown table (#15081) |
| 1000 | ee3a9fcf88fe5b5e1213711e05861b83cd4fdfe6 | ec428b02c347767f24c78111309e3f30d2ada289 | compilade | git@compilade.net | 2025-08-05T05:27:45-04:00 | GitHub | noreply@github.com | 2025-08-05T11:27:45+02:00 | | context : fix index overflow on huge outputs (#15080) |
| 1001 | ec428b02c347767f24c78111309e3f30d2ada289 | 19f68fa5a4c3bf796de52a6db9008e77d29f423a | Diego Devesa | slarengh@gmail.com | 2025-08-04T16:05:36-07:00 | GitHub | noreply@github.com | 2025-08-05T01:05:36+02:00 | | llama : add --n-cpu-moe option (#15077) |
| 1002 | 19f68fa5a4c3bf796de52a6db9008e77d29f423a | 41613437ffee0dbccad684fc744788bc504ec213 | compilade | git@compilade.net | 2025-08-04T17:26:52-04:00 | GitHub | noreply@github.com | 2025-08-04T23:26:52+02:00 | | imatrix : warn when GGUF imatrix is saved without .gguf suffix (#15076) |
| 1003 | 41613437ffee0dbccad684fc744788bc504ec213 | e5bebe5251cee2678e8531aa1598ca21b3c6ce1d | Christian Kastner | ckk@kvr.at | 2025-08-04T21:29:14+02:00 | GitHub | noreply@github.com | 2025-08-04T21:29:14+02:00 | | cmake: Add GGML_BACKEND_DIR option (#15074) |
| 1004 | e5bebe5251cee2678e8531aa1598ca21b3c6ce1d | ef0144c087b33e5b8da42d529ac71aaf05cb49df | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-04T21:01:48+02:00 | GitHub | noreply@github.com | 2025-08-04T21:01:48+02:00 | | gguf-py : add --chat-template-file to gguf_new_metadata (#15075) |
| 1005 | ef0144c087b33e5b8da42d529ac71aaf05cb49df | 2721257e3e2c4c944ac8a08221113ee7cb503f1b | Sam | sammcj@users.noreply.github.com | 2025-08-05T04:29:25+10:00 | GitHub | noreply@github.com | 2025-08-04T20:29:25+02:00 | | model: support GLM 4.5 family of models (#14939) |
| 1006 | 2721257e3e2c4c944ac8a08221113ee7cb503f1b | 587d0118f50b7e8f4bafbcdd218aefd9da0272e1 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-04T18:11:02+02:00 | GitHub | noreply@github.com | 2025-08-04T18:11:02+02:00 | | quantize : fix confusing error message if ftype is invalid (#15071) |
| 1007 | 587d0118f50b7e8f4bafbcdd218aefd9da0272e1 | 5aa1105da24a8dd1661cea3db0582c9b2c2f54d3 | Reese Levine | reeselevine1@gmail.com | 2025-08-04T08:52:43-07:00 | GitHub | noreply@github.com | 2025-08-04T08:52:43-07:00 | | ggml: WebGPU backend host improvements and style fixing (#14978) |
| 1008 | 5aa1105da24a8dd1661cea3db0582c9b2c2f54d3 | d31192b4ee1441bbbecd3cbf9e02633368bdc4f5 | Jeff Bolz | jbolz@nvidia.com | 2025-08-04T00:09:19-05:00 | GitHub | noreply@github.com | 2025-08-04T07:09:19+02:00 | | vulkan: fix build when using glslang that does not support coopmat2 (#15062) |
| 1009 | d31192b4ee1441bbbecd3cbf9e02633368bdc4f5 | 0a2f5496bef9e54e5f42d6c2c3ad9eb7b379aed0 | compilade | git@compilade.net | 2025-08-03T16:00:05-04:00 | GitHub | noreply@github.com | 2025-08-03T22:00:05+02:00 | | imatrix : use GGUF by default (#14842) |
| 1010 | 0a2f5496bef9e54e5f42d6c2c3ad9eb7b379aed0 | 11a3811164ef2d75393c6b0a632f4c608e3e3dd2 | compilade | git@compilade.net | 2025-08-03T15:49:13-04:00 | GitHub | noreply@github.com | 2025-08-03T21:49:13+02:00 | | imatrix : fix 3d activation handling for hybrid and recurrent models (#14994) |
| 1011 | 11a3811164ef2d75393c6b0a632f4c608e3e3dd2 | 97366dc6abdd0bdc74260bd3c42bd06f0feb7428 | compilade | git@compilade.net | 2025-08-03T15:43:07-04:00 | GitHub | noreply@github.com | 2025-08-03T21:43:07+02:00 | | memory : handle kv_unified for hybrid models (#15050) |
| 1012 | 97366dc6abdd0bdc74260bd3c42bd06f0feb7428 | 83bc2f288c0e08e676d9beca9c4669197e920593 | Csaba Kecskemeti | csaba.kecskemeti@gmail.com | 2025-08-03T12:38:18-07:00 | GitHub | noreply@github.com | 2025-08-03T21:38:18+02:00 | | vocab : JetBrains Mellum pre-tokenizer (#15045) |
| 1013 | 83bc2f288c0e08e676d9beca9c4669197e920593 | 6c7a441161080551ce8a52ba32563b6295067192 | Gabriel Larson | 55459720+gabriellarson@users.noreply.github.com | 2025-08-03T09:56:25-05:00 | GitHub | noreply@github.com | 2025-08-03T16:56:25+02:00 | | model : add text-only support for Kimi-VL (and find special tokens in text_config) (#15051) |
| 1014 | 6c7a441161080551ce8a52ba32563b6295067192 | 5c0eb5ef544aeefd81c303e03208f768e158d93c | Jeff Bolz | jbolz@nvidia.com | 2025-08-03T07:23:57-05:00 | GitHub | noreply@github.com | 2025-08-03T14:23:57+02:00 | | vulkan: Use coopmat2 for conv2d (#14982) |
| 1015 | 5c0eb5ef544aeefd81c303e03208f768e158d93c | 03d46982180c2fb624bd2a233e46426ab22be5d1 | lhez | lih@qti.qualcomm.com | 2025-08-02T10:51:18-07:00 | GitHub | noreply@github.com | 2025-08-02T19:51:18+02:00 | | opencl: fix adreno compiler detection logic (#15029) |
| 1016 | 03d46982180c2fb624bd2a233e46426ab22be5d1 | 3303c19b1691088275ee864a823697177c94a15d | Johannes Gäßler | johannesg@5d6.de | 2025-08-02T16:37:08+02:00 | GitHub | noreply@github.com | 2025-08-02T16:37:08+02:00 | | CUDA: use mma FA kernel for gqa > 4 on RTX 4000 (#15035) |
| 1017 | 3303c19b1691088275ee864a823697177c94a15d | 4fdea540bda4648f98b85e8ee9dc66db4bfb5945 | leejet | leejet714@gmail.com | 2025-08-02T22:15:36+08:00 | GitHub | noreply@github.com | 2025-08-02T17:15:36+03:00 | | cuda: make im2col a little faster (#15025) |
| 1018 | 4fdea540bda4648f98b85e8ee9dc66db4bfb5945 | a4569c41fd2253c89ef52fc2378687bdbf42f61a | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-08-02T16:14:57+02:00 | GitHub | noreply@github.com | 2025-08-02T17:14:57+03:00 | | kv-cache : skip alignment of n_stream in kv-cache log msg [no ci] (#15040) |
| 1019 | a4569c41fd2253c89ef52fc2378687bdbf42f61a | 15e92fd33791e60a4ddb5970b47242a855c27117 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-02T17:14:21+03:00 | GitHub | noreply@github.com | 2025-08-02T17:14:21+03:00 | | llama : enable LLAMA_SET_ROWS=1 by default (#14959) |
| 1020 | 15e92fd33791e60a4ddb5970b47242a855c27117 | 2bf3fbf0b54f97aef2b388b76d222789e1c170f1 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-02T17:13:05+03:00 | GitHub | noreply@github.com | 2025-08-02T17:13:05+03:00 | | cuda, sycl : fix batched gemm when ne02 == 1 && ne03 > 1 (#15038) |
| 1021 | 2bf3fbf0b54f97aef2b388b76d222789e1c170f1 | 711d5e6fe66eb6cd7a10d71cec4567321848be08 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-08-02T14:39:01+02:00 | GitHub | noreply@github.com | 2025-08-02T14:39:01+02:00 | | ci : check that pre-tokenizer hashes are up-to-date (#15032) |
| 1022 | 711d5e6fe66eb6cd7a10d71cec4567321848be08 | f738989dcb9ccbe468c945553eafbeef7b869675 | Douglas Hanley | thesecretaryofwar@gmail.com | 2025-08-02T05:51:02-05:00 | GitHub | noreply@github.com | 2025-08-02T12:51:02+02:00 | | convert : fix Qwen3-Embedding pre-tokenizer hash (#15030) |
| 1023 | f738989dcb9ccbe468c945553eafbeef7b869675 | 4cb208c93c1c938591a5b40354e2a6f9b94489bc | Jhen-Jie Hong | iainst0409@gmail.com | 2025-08-02T18:04:48+08:00 | GitHub | noreply@github.com | 2025-08-02T18:04:48+08:00 | | chat : fix multiple tool_calls on hermes-2-pro (#14962) |
| 1024 | 4cb208c93c1c938591a5b40354e2a6f9b94489bc | 3025b621d12a6931ff5e9775d4f644719980ad91 | Jeff Bolz | jbolz@nvidia.com | 2025-08-02T04:21:37-05:00 | GitHub | noreply@github.com | 2025-08-02T11:21:37+02:00 | | vulkan: coopmat2 mul_mat optimizations (#14934) |
| 1025 | 3025b621d12a6931ff5e9775d4f644719980ad91 | ec0b18802c91badd3ff1388ffd09ee163251bd72 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-02T17:20:40+08:00 | GitHub | noreply@github.com | 2025-08-02T17:20:40+08:00 | | llama-bench: rename DB table name from test to llama_bench (#15003) |
| 1026 | ec0b18802c91badd3ff1388ffd09ee163251bd72 | 339bd0268c498c89529cd0e90c44883c211e3745 | Jeff Bolz | jbolz@nvidia.com | 2025-08-02T03:48:30-05:00 | GitHub | noreply@github.com | 2025-08-02T10:48:30+02:00 | | vulkan: Support ne[3]>1 in noncontig matrix-vector multiply (#15015) |
| 1027 | 339bd0268c498c89529cd0e90c44883c211e3745 | f906275537d14c8fc7c6976d944233771fd6672c | Douglas Hanley | thesecretaryofwar@gmail.com | 2025-08-02T03:44:50-05:00 | GitHub | noreply@github.com | 2025-08-02T10:44:50+02:00 | | model : support Qwen3-Embedding (#15023) |
| 1028 | f906275537d14c8fc7c6976d944233771fd6672c | a9f7541ec25c4c8547daf5ff48700ad2836e2b7d | Johannes Gäßler | johannesg@5d6.de | 2025-08-02T10:12:41+02:00 | GitHub | noreply@github.com | 2025-08-02T10:12:41+02:00 | | server: enable token array inputs for OAI API (#15001) |
| 1029 | a9f7541ec25c4c8547daf5ff48700ad2836e2b7d | 9c35706b98ea271858acef4194f526a71b24cdc9 | Jeff Bolz | jbolz@nvidia.com | 2025-08-02T02:57:04-05:00 | GitHub | noreply@github.com | 2025-08-02T09:57:04+02:00 | | vulkan: optimizations for direct convolution (#14933) |
| 1030 | 9c35706b98ea271858acef4194f526a71b24cdc9 | c76b420e4ce06f7b7cdfbb0b85d02c90e5cc5a3a | Johannes Gäßler | johannesg@5d6.de | 2025-08-01T20:47:32+02:00 | GitHub | noreply@github.com | 2025-08-01T20:47:32+02:00 | | CUDA: fix MMQ nwarps for AMD with warp_size==32 (#15014) |
| 1031 | c76b420e4ce06f7b7cdfbb0b85d02c90e5cc5a3a | 0f5ccd6fd1a1f709010312933db0316867cc30b6 | l-austenfeld | 53152202+l-austenfeld@users.noreply.github.com | 2025-08-01T16:59:06+02:00 | GitHub | noreply@github.com | 2025-08-01T16:59:06+02:00 | | vendor : update vendored copy of google/minja (#15011) |
| 1032 | 0f5ccd6fd1a1f709010312933db0316867cc30b6 | 1c872f71fb8a25589efa3ee9b6bf8b517cb8caa4 | stevenkuang | stevenkuang@tencent.com | 2025-08-01T21:31:12+08:00 | GitHub | noreply@github.com | 2025-08-01T15:31:12+02:00 | | model : add hunyuan dense (#14878) |
| 1033 | 1c872f71fb8a25589efa3ee9b6bf8b517cb8caa4 | baad94885df512bb24ab01e2b22d1998fce4d00e | lhez | quic_lih@quicinc.com | 2025-08-01T04:15:44-07:00 | GitHub | noreply@github.com | 2025-08-01T13:15:44+02:00 | | opencl: add f16 for `add`, `sub`, `mul`, `div` (#14984) |
| 1034 | baad94885df512bb24ab01e2b22d1998fce4d00e | ba42794c9ead96ad52311ba1b23eefcbf3d6f63d | Srihari-mcw | 96763064+Srihari-mcw@users.noreply.github.com | 2025-08-01T11:50:33+05:30 | GitHub | noreply@github.com | 2025-08-01T09:20:33+03:00 | | ggml : Q2k interleaving implementation - x86/x64 SIMD (#14373) |
| 1035 | ba42794c9ead96ad52311ba1b23eefcbf3d6f63d | 2860d479b456e1caa026b40b829d5b13c42a8ed7 | Georgi Gerganov | ggerganov@gmail.com | 2025-08-01T06:38:12+03:00 | GitHub | noreply@github.com | 2025-08-01T06:38:12+03:00 | | graph : fix equal_seq() check (#14986) |
| 1036 | 2860d479b456e1caa026b40b829d5b13c42a8ed7 | 484b2091ce5017901483b5204c07878f171d1441 | diannao | 55k@outlook.com | 2025-08-01T10:02:34+08:00 | GitHub | noreply@github.com | 2025-08-01T10:02:34+08:00 | | docker : add cann build pipline (#14591) |
| 1037 | 484b2091ce5017901483b5204c07878f171d1441 | daf2dd788066b8b239cb7f68210e090c2124c199 | R0CKSTAR | yeahdongcn@gmail.com | 2025-08-01T08:47:27+08:00 | GitHub | noreply@github.com | 2025-08-01T08:47:27+08:00 | | compare-commits.sh: support both llama-bench and test-backend-ops (#14392) |
| 1038 | daf2dd788066b8b239cb7f68210e090c2124c199 | a06ed5feaec9f935fbf662035b2673167bc88460 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-31T20:32:18+01:00 | GitHub | noreply@github.com | 2025-07-31T21:32:18+02:00 | | quantize : skip tensor override when in fallback mode (#14995) |
| 1039 | a06ed5feaec9f935fbf662035b2673167bc88460 | 784524053d986255e383dae45235e4f76c6792a7 | Diego Devesa | slarengh@gmail.com | 2025-07-31T11:15:41-07:00 | GitHub | noreply@github.com | 2025-07-31T20:15:41+02:00 | | llama : add simple option to enable CPU for MoE weights (--cpu-moe) (#14992) |
| 1040 | 784524053d986255e383dae45235e4f76c6792a7 | d6818d06a6237631523bc0f45d42e79482667948 | Aman Gupta | amangupta052@gmail.com | 2025-08-01T01:22:58+08:00 | GitHub | noreply@github.com | 2025-08-01T01:22:58+08:00 | | Fix params bug in diffusion example (#14993) |
| 1041 | d6818d06a6237631523bc0f45d42e79482667948 | e08a98826bcce70a3377592c51d0efed99eafe07 | Diego Devesa | slarengh@gmail.com | 2025-07-31T09:11:34-07:00 | GitHub | noreply@github.com | 2025-07-31T18:11:34+02:00 | | llama : allow other bufts when overriding to CPU, add --no-repack option (#14990) |
| 1042 | e08a98826bcce70a3377592c51d0efed99eafe07 | 952a47f455fbd92e2659b98b9b6317a2dafeb532 | Ruben Ortlam | picard12@live.de | 2025-07-31T17:46:54+02:00 | GitHub | noreply@github.com | 2025-07-31T17:46:54+02:00 | | Vulkan: Fix minor debug mode issues (#14899) |
| 1043 | 952a47f455fbd92e2659b98b9b6317a2dafeb532 | 36e5fe7bcd6640f4ff974d4c5cb04a13b5b29bce | tc-mb | 157115220+tc-mb@users.noreply.github.com | 2025-07-31T23:22:17+08:00 | GitHub | noreply@github.com | 2025-07-31T17:22:17+02:00 | | mtmd : support MiniCPM-V 4.0 (#14983) |
| 1044 | 36e5fe7bcd6640f4ff974d4c5cb04a13b5b29bce | 94933c8c2eeaa9a7983e3f6c08af76bd86724094 | Csaba Kecskemeti | csaba.kecskemeti@gmail.com | 2025-07-31T07:59:49-07:00 | GitHub | noreply@github.com | 2025-07-31T10:59:49-04:00 | | MODEL_TENSOR.SSM_DT_NORM has defined twice (#14991) |
| 1045 | 94933c8c2eeaa9a7983e3f6c08af76bd86724094 | c1dacaa99b4ead6edbac928cd2da59436573f6b0 | g2mt | 166577174+g2mt@users.noreply.github.com | 2025-07-31T05:25:23-07:00 | GitHub | noreply@github.com | 2025-07-31T14:25:23+02:00 | | server : implement universal assisted decoding (#12635) |
| 1046 | c1dacaa99b4ead6edbac928cd2da59436573f6b0 | a9f77a8be348deeb11fb3d54d412bf583003c90d | Dongliang Wei | 121270393+wdl339@users.noreply.github.com | 2025-07-31T20:12:20+08:00 | GitHub | noreply@github.com | 2025-07-31T14:12:20+02:00 | | llama : merge build_moe_ffn_from_probs function into build_moe_ffn (#14968) |
| 1047 | a9f77a8be348deeb11fb3d54d412bf583003c90d | 8a4a85627702b569d7d2810f2de06a4321656e9d | Lukas Straub | lukasstraub2@web.de | 2025-07-31T14:08:23+02:00 | GitHub | noreply@github.com | 2025-07-31T14:08:23+02:00 | | server : add openai-style logit_bias support (#14946) |
| 1048 | 8a4a85627702b569d7d2810f2de06a4321656e9d | 11490b36723d511d75fb601995c79b5c363ba3a2 | Aman Gupta | amangupta052@gmail.com | 2025-07-31T19:49:09+08:00 | GitHub | noreply@github.com | 2025-07-31T19:49:09+08:00 | | Add LLaDA 8b Diffusion model (#14771) |
| 1049 | 11490b36723d511d75fb601995c79b5c363ba3a2 | 66625a59a54d0a7504eda4c4e83abfcd83ba1cf8 | hipudding | huafengchun@gmail.com | 2025-07-31T19:47:20+08:00 | GitHub | noreply@github.com | 2025-07-31T19:47:20+08:00 | | CANN: Improve loading efficiency after converting weights to NZ format. (#14985) |
| 1050 | 66625a59a54d0a7504eda4c4e83abfcd83ba1cf8 | 6e6725459a892b49602b596339de4916c7c7965a | compilade | git@compilade.net | 2025-07-31T01:02:46-04:00 | GitHub | noreply@github.com | 2025-07-31T08:02:46+03:00 | | graph : reduce splits for recurrent and hybrid models (#14825) |
| 1051 | 6e6725459a892b49602b596339de4916c7c7965a | e9192bec564780bd4313ad6524d20a0ab92797db | lhez | lih@qti.qualcomm.com | 2025-07-30T14:56:55-07:00 | GitHub | noreply@github.com | 2025-07-30T14:56:55-07:00 | | opencl: add `mul_mat_f32_f32_l4_lm` and `mul_mat_f16_f32_l4_lm` (#14809) |
| 1052 | e9192bec564780bd4313ad6524d20a0ab92797db | 41e78c567e9a8c652e405f4f909deb598deecd31 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-30T20:11:56+01:00 | GitHub | noreply@github.com | 2025-07-30T21:11:56+02:00 | | quantize : fix using combined imatrix GGUFs (multiple datasets) (#14973) |
| 1053 | 41e78c567e9a8c652e405f4f909deb598deecd31 | ad4a700117d1746799d2d6599e526e2c3a7938d2 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-07-30T18:07:11+02:00 | GitHub | noreply@github.com | 2025-07-30T18:07:11+02:00 | | server : add support for `embd_normalize` parameter (#14964) |
| 1054 | ad4a700117d1746799d2d6599e526e2c3a7938d2 | e32a4ec60ee472acf3fe82b5966976fd6965ac6b | uvos | carl@uvos.xyz | 2025-07-30T17:38:06+02:00 | GitHub | noreply@github.com | 2025-07-30T17:38:06+02:00 | | HIP: enable mfma mmq on gfx908 and gfx90a for select datatypes and shapes (#14949) |
| 1055 | e32a4ec60ee472acf3fe82b5966976fd6965ac6b | e228de94496636681b36690247d96f97d5f76c0d | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T16:03:13+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T17:33:11+03:00 | | sync : ggml |
| 1056 | e228de94496636681b36690247d96f97d5f76c0d | 73a8e5ca0372f4dcaf1aed4e42261723da0914aa | Kai Pastor | dg0yt@darc.de | 2025-07-30T14:53:16+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T17:33:11+03:00 | | cmake : Fix BLAS link interface (ggml/1316) |
| 1057 | 73a8e5ca0372f4dcaf1aed4e42261723da0914aa | 92b8810ec7aa6d778bc287cc918443cf67b962e2 | Kai Pastor | dg0yt@darc.de | 2025-07-30T14:52:26+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T17:33:11+03:00 | | vulkan : fix 32-bit builds (ggml/1313) |
| 1058 | 92b8810ec7aa6d778bc287cc918443cf67b962e2 | 00131d6eaf4df029e1ec84de868c2c5957503007 | Johannes Gäßler | johannesg@5d6.de | 2025-07-30T15:46:13+02:00 | GitHub | noreply@github.com | 2025-07-30T15:46:13+02:00 | | CUDA: skip masked KV slices for all FA kernels (#14924) |
| 1059 | 00131d6eaf4df029e1ec84de868c2c5957503007 | 1e15bfd42c3938506e0da5939bf7f42780965f01 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T15:12:02+03:00 | GitHub | noreply@github.com | 2025-07-30T15:12:02+03:00 | | tests : update for LLAMA_SET_ROWS=1 (#14961) |
| 1060 | 1e15bfd42c3938506e0da5939bf7f42780965f01 | a118d80233d3bf92569c051346fd2638f87bf202 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-30T13:52:11+03:00 | GitHub | noreply@github.com | 2025-07-30T13:52:11+03:00 | | graph : fix stack-use-after-return (#14960) |
| 1061 | a118d80233d3bf92569c051346fd2638f87bf202 | 61550f8231dc0aa478e2f537c5009ede6878ce22 | Douglas Hanley | thesecretaryofwar@gmail.com | 2025-07-30T00:25:05-05:00 | GitHub | noreply@github.com | 2025-07-30T08:25:05+03:00 | | embeddings: fix extraction of CLS pooling results (#14927) |
| 1062 | 61550f8231dc0aa478e2f537c5009ede6878ce22 | aa79524c51fb014f8df17069d31d7c44b9ea6cb8 | Xinpeng Dou | 15529241576@163.com | 2025-07-30T08:39:24+08:00 | GitHub | noreply@github.com | 2025-07-30T08:39:24+08:00 | | CANN: update ops docs (#14935) |
| 1063 | aa79524c51fb014f8df17069d31d7c44b9ea6cb8 | b77d11179d7efe68ce5c913006f2ef9ee17cc8f7 | uvos | carl@uvos.xyz | 2025-07-29T20:23:04+02:00 | GitHub | noreply@github.com | 2025-07-29T20:23:04+02:00 | | HIP: remove the use of __HIP_PLATFORM_AMD__, explicitly support only AMD targets (#14945) |
| 1064 | b77d11179d7efe68ce5c913006f2ef9ee17cc8f7 | c7aa1364fd59b2ac06fd9e0a719253d968472dc3 | uvos | carl@uvos.xyz | 2025-07-29T17:44:30+02:00 | GitHub | noreply@github.com | 2025-07-29T17:44:30+02:00 | | HIP: add GGML_HIP_MMQ_MFMA option to allow disableing the MFMA path. (#14930) |
| 1065 | c7aa1364fd59b2ac06fd9e0a719253d968472dc3 | 1a67fcc30677e96dda76bb1b290788e7d8852b51 | uvos | carl@uvos.xyz | 2025-07-29T17:43:43+02:00 | GitHub | noreply@github.com | 2025-07-29T17:43:43+02:00 | | HIP: Ignore unsupported unroll transformation in fattn-vec (#14931) |
| 1066 | 1a67fcc30677e96dda76bb1b290788e7d8852b51 | 204f2cf168dc01ca7b200b1510e0ff585ca9a92e | kallewoof | karljohan-alm@garage.co.jp | 2025-07-30T00:05:38+09:00 | GitHub | noreply@github.com | 2025-07-29T17:05:38+02:00 | | common : avoid logging partial messages (which can contain broken UTF-8 sequences) (#14937) |
| 1067 | 204f2cf168dc01ca7b200b1510e0ff585ca9a92e | 138b288b594eeb19d58e0ab0d74eae32009c9c20 | hipudding | huafengchun@gmail.com | 2025-07-29T22:36:43+08:00 | GitHub | noreply@github.com | 2025-07-29T22:36:43+08:00 | | CANN: Add ggml_set_rows (#14943) |
| 1068 | 138b288b594eeb19d58e0ab0d74eae32009c9c20 | bbd0f917797e9d524680f1b30d34a46eb06d7651 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-29T14:22:03+02:00 | GitHub | noreply@github.com | 2025-07-29T14:22:03+02:00 | | cuda : add softcap fusion (#14907) |
| 1069 | bbd0f917797e9d524680f1b30d34a46eb06d7651 | 0a5036bee9cfb946870689db4400e9e0d17844c9 | Johannes Gäßler | johannesg@5d6.de | 2025-07-29T10:40:50+02:00 | GitHub | noreply@github.com | 2025-07-29T10:40:50+02:00 | | server-bench: make seed choice configurable (#14929) |
| 1070 | 0a5036bee9cfb946870689db4400e9e0d17844c9 | 8ad7b3e65b5834e5574c2f5640056c9047b5d93b | Aman Gupta | amangupta052@gmail.com | 2025-07-29T14:45:18+08:00 | GitHub | noreply@github.com | 2025-07-29T14:45:18+08:00 | | CUDA: add roll (#14919) |
| 1071 | 8ad7b3e65b5834e5574c2f5640056c9047b5d93b | bda62193b2a6bebbf515c3c389303094a44458c1 | lhez | lih@qti.qualcomm.com | 2025-07-28T09:50:17-07:00 | GitHub | noreply@github.com | 2025-07-28T18:50:17+02:00 | | opencl : add ops docs (#14910) |
| 1072 | bda62193b2a6bebbf515c3c389303094a44458c1 | c556418b600ad5792440942079d93e393595688b | Leonard Mosescu | tlemo@users.noreply.github.com | 2025-07-28T09:04:27-07:00 | GitHub | noreply@github.com | 2025-07-28T18:04:27+02:00 | | test-backend-ops : extend test case filtering (#14865) |
| 1073 | c556418b600ad5792440942079d93e393595688b | db16e2831c0f344f041af3d067db81c42b16eb22 | Radoslav Gerganov | rgerganov@gmail.com | 2025-07-28T18:59:04+03:00 | GitHub | noreply@github.com | 2025-07-28T18:59:04+03:00 | | llama-bench : use local GPUs along with RPC servers (#14917) |
| 1074 | db16e2831c0f344f041af3d067db81c42b16eb22 | cd1fce6d4f9c191f1c7429cc96f61281c3b63ffc | xctan | xc-tan@outlook.com | 2025-07-28T23:40:24+08:00 | GitHub | noreply@github.com | 2025-07-28T17:40:24+02:00 | | ggml-cpu : deduplicate scalar implementations (#14897) |
| 1075 | cd1fce6d4f9c191f1c7429cc96f61281c3b63ffc | 00fa15fedc79263fa0285e6a3bbb0cfb3e3878a2 | Akarshan Biswas | akarshan@menlo.ai | 2025-07-28T20:32:15+05:30 | GitHub | noreply@github.com | 2025-07-28T20:32:15+05:30 | | SYCL: Add set_rows support for quantized types (#14883) |
| 1076 | 00fa15fedc79263fa0285e6a3bbb0cfb3e3878a2 | 946b1f685909c8c9c044f145bce819c02f327eaa | Xuan-Son Nguyen | son@huggingface.co | 2025-07-28T15:01:48+02:00 | GitHub | noreply@github.com | 2025-07-28T15:01:48+02:00 | | mtmd : add support for Voxtral (#14862) |
| 1077 | 946b1f685909c8c9c044f145bce819c02f327eaa | 6c6e397affc4fac717e718364fb4b635cec6433a | Johannes Gäßler | johannesg@5d6.de | 2025-07-28T14:30:22+02:00 | GitHub | noreply@github.com | 2025-07-28T14:30:22+02:00 | | CUDA: fix pointer incrementation in FA (#14916) |
| 1078 | 6c6e397affc4fac717e718364fb4b635cec6433a | afc0e8969896ada62238da07b98731e5a4b12ba4 | Dongliang Wei | 121270393+wdl339@users.noreply.github.com | 2025-07-28T19:47:00+08:00 | GitHub | noreply@github.com | 2025-07-28T13:47:00+02:00 | | model : add support for SmallThinker series (#14898) |
| 1079 | afc0e8969896ada62238da07b98731e5a4b12ba4 | a5771c9eea801f573dc7375416e4c3306543563d | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-07-28T11:05:53+01:00 | GitHub | noreply@github.com | 2025-07-28T11:05:53+01:00 | | sycl: refactor quantization to q8_1 (#14815) |
| 1080 | a5771c9eea801f573dc7375416e4c3306543563d | c35f9eaf095e0db3aa9b77b56846bab5d3bf5661 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-28T11:01:03+03:00 | GitHub | noreply@github.com | 2025-07-28T10:01:03+02:00 | | ops : update BLAS (#14914) |
| 1081 | c35f9eaf095e0db3aa9b77b56846bab5d3bf5661 | 1f45f2890ef7f365ba0a45e08a8d1f46b8bc6b9e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-28T08:22:56+03:00 | GitHub | noreply@github.com | 2025-07-28T08:22:56+03:00 | | ops : update Metal (#14912) |
| 1082 | 1f45f2890ef7f365ba0a45e08a8d1f46b8bc6b9e | 613c5095c33de9b7bbf1097d7c32510f51d58b01 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-28T08:14:20+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-28T08:15:01+03:00 | | sync : ggml |
| 1083 | 613c5095c33de9b7bbf1097d7c32510f51d58b01 | 7f97599581fcf0c37432dd3b1f503b91bed97695 | Kai Pastor | dg0yt@darc.de | 2025-07-24T19:58:02+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-28T08:15:01+03:00 | | cmake : Indent ggml-config.cmake (ggml/1310) |
| 1084 | 7f97599581fcf0c37432dd3b1f503b91bed97695 | bf78f5439ee8e82e367674043303ebf8e92b4805 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-27T22:31:11+01:00 | GitHub | noreply@github.com | 2025-07-27T23:31:11+02:00 | | quantize : update README.md (#14905) |
| 1085 | bf78f5439ee8e82e367674043303ebf8e92b4805 | bbfc84927481d1e59d8af6939b93b76850f0ab53 | Ruben Ortlam | picard12@live.de | 2025-07-27T15:33:08+02:00 | GitHub | noreply@github.com | 2025-07-27T15:33:08+02:00 | | vulkan: add ops docs (#14900) |
| 1086 | bbfc84927481d1e59d8af6939b93b76850f0ab53 | ca0ef2dddb022cb1337d775cd05cd27d7808aff4 | Akarshan Biswas | akarshan@menlo.ai | 2025-07-27T17:52:58+05:30 | GitHub | noreply@github.com | 2025-07-27T17:52:58+05:30 | | SYCL: add ops doc (#14901) |
| 1087 | ca0ef2dddb022cb1337d775cd05cd27d7808aff4 | 89d1029559bd2968f76db854f9f113d73e34527c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-07-27T12:10:51+02:00 | GitHub | noreply@github.com | 2025-07-27T12:10:51+02:00 | | llama : clarify comment about pp and tg graphs [no ci] (#14895) |
| 1088 | 89d1029559bd2968f76db854f9f113d73e34527c | f1a4e72de5950ab7136aeadddd675caf30dd6b3f | Erik Scholz | Green-Sky@users.noreply.github.com | 2025-07-27T12:04:33+02:00 | GitHub | noreply@github.com | 2025-07-27T12:04:33+02:00 | | vulkan : add fp16 support for the conv_2d kernel (#14872) |
| 1089 | f1a4e72de5950ab7136aeadddd675caf30dd6b3f | 4762ad7316dcdec20016ab5985fb46a27902204d | Jeff Bolz | jbolz@nvidia.com | 2025-07-27T04:05:34-05:00 | GitHub | noreply@github.com | 2025-07-27T11:05:34+02:00 | | vulkan: skip empty set_rows to avoid invalid API usage (#14860) |
| 1090 | 4762ad7316dcdec20016ab5985fb46a27902204d | 1dc9614e0673e794d2e2bf88ba04f7d57b63a57b | Gabriel Larson | 55459720+gabriellarson@users.noreply.github.com | 2025-07-27T03:18:37-05:00 | GitHub | noreply@github.com | 2025-07-27T11:18:37+03:00 | | model : make rope_yarn_log_mul optional for deepseek2 (#14896) |
| 1091 | 1dc9614e0673e794d2e2bf88ba04f7d57b63a57b | 446595b9b3a113d9ba10506922c3a156cca9d477 | Shunta Saito | shunta.saito@gmail.com | 2025-07-27T16:38:44+09:00 | GitHub | noreply@github.com | 2025-07-27T09:38:44+02:00 | | llama : fix kq_scale for the attention layers of PLaMo2 (#14892) |
| 1092 | 446595b9b3a113d9ba10506922c3a156cca9d477 | 66906cd82a4a1fd10151707cee3f66cb61fc4055 | Aman Gupta | amangupta052@gmail.com | 2025-07-27T09:36:43+08:00 | GitHub | noreply@github.com | 2025-07-27T09:36:43+08:00 | | Docs: add instructions for adding backends (#14889) |
| 1093 | 66906cd82a4a1fd10151707cee3f66cb61fc4055 | 11dd5a44eb180e1d69fac24d3852b5222d66fb7f | deepsek | 166548550+deepsek@users.noreply.github.com | 2025-07-26T18:28:14-04:00 | GitHub | noreply@github.com | 2025-07-27T00:28:14+02:00 | | HIP: Enable Matrix cores for MMQ Kernels, Enable stream-K for CDNA 3 (#14624) |
| 1094 | 11dd5a44eb180e1d69fac24d3852b5222d66fb7f | 9b8f3c6c776d77045ee4f7f13fdf7863dfe59e5b | hipudding | huafengchun@gmail.com | 2025-07-26T17:56:18+08:00 | GitHub | noreply@github.com | 2025-07-26T17:56:18+08:00 | | CANN: Implement GLU ops (#14884) |
| 1095 | 9b8f3c6c776d77045ee4f7f13fdf7863dfe59e5b | c7f3169cd523140a288095f2d79befb20a0b73f4 | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-26T10:36:02+08:00 | GitHub | noreply@github.com | 2025-07-26T10:36:02+08:00 | | musa: fix build warnings (unused variable) (#14869) |
| 1096 | c7f3169cd523140a288095f2d79befb20a0b73f4 | 793c0d7f46384001738c337d7afa46b45ae32745 | Aaron Teo | aaron.teo1@ibm.com | 2025-07-26T01:09:03+08:00 | GitHub | noreply@github.com | 2025-07-25T19:09:03+02:00 | | ggml-cpu : disable GGML_NNPA by default due to instability (#14880) |
| 1097 | 793c0d7f46384001738c337d7afa46b45ae32745 | ce111d39d666f4a3b6a561ad020f6feb8cc67790 | Gabe Goodhart | ghart@us.ibm.com | 2025-07-25T10:47:39-06:00 | GitHub | noreply@github.com | 2025-07-25T10:47:39-06:00 | | metal: SSM_SCAN performance (#14743) |
| 1098 | ce111d39d666f4a3b6a561ad020f6feb8cc67790 | e7fecba93416c8aaf343932b5e36dd5e208a815e | lhez | lih@qti.qualcomm.com | 2025-07-25T08:12:13-07:00 | GitHub | noreply@github.com | 2025-07-25T17:12:13+02:00 | | opencl: add fused `rms_norm_mul` (#14841) |
| 1099 | e7fecba93416c8aaf343932b5e36dd5e208a815e | e2b7621e7c265a6739225125cf9c534f471b3472 | wooksong | wook16.song@samsung.com | 2025-07-25T23:25:05+09:00 | GitHub | noreply@github.com | 2025-07-25T16:25:05+02:00 | | docs : update HOWTO‑add‑model.md for ModelBase and new model classes (#14874) |
| 1100 | e2b7621e7c265a6739225125cf9c534f471b3472 | c1dbea752a630169c104118fd0c82f3ffcb19c91 | Oliver Simons | osimons@nvidia.com | 2025-07-25T13:29:57+02:00 | GitHub | noreply@github.com | 2025-07-25T14:29:57+03:00 | | ggml : remove invalid portPos specifiers from dot files (#14838) |
| 1101 | c1dbea752a630169c104118fd0c82f3ffcb19c91 | 749e0d27f0247337869f4698f59dd7fafba94326 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-25T14:28:06+03:00 | GitHub | noreply@github.com | 2025-07-25T14:28:06+03:00 | | context : restore preemptive sched reset when LLAMA_SET_ROWS=0 (#14870) |
| 1102 | 749e0d27f0247337869f4698f59dd7fafba94326 | 64bf1c3744053cf7def10aeed21ff48883ee755b | kiwi | 122582483+kiwi142857@users.noreply.github.com | 2025-07-25T19:08:04+08:00 | GitHub | noreply@github.com | 2025-07-25T13:08:04+02:00 | | mtmd : fix 32-bit narrowing issue in export-lora and mtmd clip (#14503) |
| 1103 | 64bf1c3744053cf7def10aeed21ff48883ee755b | c12bbde37258c6d053322d1cedd4a9672109ae58 | Chris Rohlf | chris.rohlf@gmail.com | 2025-07-25T06:17:02-04:00 | GitHub | noreply@github.com | 2025-07-25T12:17:02+02:00 | | rpc : check for null buffers in get/set/copy tensor endpoints (#14868) |
| 1104 | c12bbde37258c6d053322d1cedd4a9672109ae58 | 3f4fc97f1d745f1d5d3c853949503136d419e6de | Diego Devesa | slarengh@gmail.com | 2025-07-25T01:07:26-07:00 | GitHub | noreply@github.com | 2025-07-25T11:07:26+03:00 | | sched : fix multiple evaluations of the same graph with pipeline parallelism (#14855) |
| 1105 | 3f4fc97f1d745f1d5d3c853949503136d419e6de | 2df255da3cea108de0ae9b302ffdd31674b6d88d | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-25T03:05:37+08:00 | GitHub | noreply@github.com | 2025-07-24T20:05:37+01:00 | | musa: upgrade musa sdk to rc4.2.0 (#14498) |
| 1106 | 2df255da3cea108de0ae9b302ffdd31674b6d88d | 60f816a79dd74007158745530e71738aa6caa67e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T18:30:33+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T20:27:23+03:00 | | sync : ggml |
| 1107 | 60f816a79dd74007158745530e71738aa6caa67e | 5592f278b6ec6aa4a1793e89e8c61838a12ebc9d | Kai Pastor | dg0yt@darc.de | 2025-07-22T20:13:21+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T20:27:23+03:00 | | cmake : fix usage issues (ggml/1257) |
| 1108 | 5592f278b6ec6aa4a1793e89e8c61838a12ebc9d | e4868d16d24dec55e61bcaadaca28feed8f98b13 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-07-21T15:53:12+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T20:27:23+03:00 | | ggml-cpu : remove stdlib include from repack.cpp (ggml/1276) |
| 1109 | e4868d16d24dec55e61bcaadaca28feed8f98b13 | 820de57d4faa427a3d0bfb14e48057247fae036e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T16:31:48+03:00 | GitHub | noreply@github.com | 2025-07-24T16:31:48+03:00 | | context : perform output reorder lazily upon access after sync (#14853) |
| 1110 | 820de57d4faa427a3d0bfb14e48057247fae036e | cb4a63aad6650c2b536a7578403935388cb2920e | Xuan-Son Nguyen | son@huggingface.co | 2025-07-24T13:59:56+02:00 | GitHub | noreply@github.com | 2025-07-24T13:59:56+02:00 | | chat : fix kimi-k2 chat template (#14852) |
| 1111 | cb4a63aad6650c2b536a7578403935388cb2920e | 86f5623d904cfd392fdeb14a143097b4074660f6 | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-07-24T11:09:57+01:00 | GitHub | noreply@github.com | 2025-07-24T11:09:57+01:00 | | sycl: fixed semantics of block offset calculation (#14814) |
| 1112 | 86f5623d904cfd392fdeb14a143097b4074660f6 | 39cffdf18855e0d2beba62572542251d87421e73 | yummy | 57988893+jk3456a@users.noreply.github.com | 2025-07-24T17:50:51+08:00 | GitHub | noreply@github.com | 2025-07-24T11:50:51+02:00 | | llama : fix MiniCPM inference after Granite Four changes (#14850) |
| 1113 | 39cffdf18855e0d2beba62572542251d87421e73 | 065908cb09adff8b2a5f1173f80050d5d9a6790b | Pouya | PooyaGhahramanian@Gmail.com | 2025-07-24T12:26:44+03:00 | GitHub | noreply@github.com | 2025-07-24T11:26:44+02:00 | | docs: add libcurl-dev install hint for Linux distros (#14801) |
| 1114 | 065908cb09adff8b2a5f1173f80050d5d9a6790b | 4ec6291a2407404de52239c1f9ca66c07e7fb28b | Georgi Gerganov | ggerganov@gmail.com | 2025-07-24T10:24:05+03:00 | GitHub | noreply@github.com | 2025-07-24T10:24:05+03:00 | | metal : fix fusion across different encoders (#14849) |
| 1115 | 4ec6291a2407404de52239c1f9ca66c07e7fb28b | a12363bbf0ca11f787dee9756861043136edc0df | Donghyeon Jeong | 54725479+djeong20@users.noreply.github.com | 2025-07-24T13:50:41+09:00 | GitHub | noreply@github.com | 2025-07-24T12:50:41+08:00 | | sycl: fix undefined variable in work group size check (#14843) |
| 1116 | a12363bbf0ca11f787dee9756861043136edc0df | a86f52b2859dae4db5a7a0bbc0f1ad9de6b43ec6 | jacekpoplawski | 67507230+jacekpoplawski@users.noreply.github.com | 2025-07-23T23:23:57+02:00 | GitHub | noreply@github.com | 2025-07-23T23:23:57+02:00 | | convert : text-only support for GLM-4.1V-9B-Thinking (#14823) |
| 1117 | a86f52b2859dae4db5a7a0bbc0f1ad9de6b43ec6 | b284197df426fb189cdcfe56a43c863a788ac756 | Johannes Gäßler | johannesg@5d6.de | 2025-07-23T21:43:25+02:00 | GitHub | noreply@github.com | 2025-07-23T21:43:25+02:00 | | CUDA: fix overflow in FA, tune performance (#14840) |
| 1118 | b284197df426fb189cdcfe56a43c863a788ac756 | 221c0e0c5841a814e95b9bfd549de9d6ae00ac6e | Johannes Gäßler | johannesg@5d6.de | 2025-07-23T18:22:30+02:00 | GitHub | noreply@github.com | 2025-07-23T18:22:30+02:00 | | CUDA: fix compilation with GGML_CUDA_F16 (#14837) |
| 1119 | 221c0e0c5841a814e95b9bfd549de9d6ae00ac6e | 07a19e27a26f76d34be62da53807f93131fb3cab | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-23T14:27:54+02:00 | GitHub | noreply@github.com | 2025-07-23T14:27:54+02:00 | | ci : correct label refactor->refactoring (#14832) |
| 1120 | 07a19e27a26f76d34be62da53807f93131fb3cab | 18f3b5ff9e5eda4e7d04bceff8ffdccb0a696ed8 | Johannes Gäßler | johannesg@5d6.de | 2025-07-23T12:35:53+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-23T14:08:09+03:00 | | CUDA: fix quantized KV cache + multiple sequences (#14822) |
| 1121 | 18f3b5ff9e5eda4e7d04bceff8ffdccb0a696ed8 | 7233358d29df3dfbf80693f2de7e5865401ad724 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T13:36:27+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-23T14:08:09+03:00 | | tests : add non-cont K,V FA tests |
| 1122 | 7233358d29df3dfbf80693f2de7e5865401ad724 | 6c88b3bb2509d980e6a64c50fdf8dd304929f770 | l3utterfly | gc.pthzfoldr@gmail.com | 2025-07-23T16:16:41+08:00 | GitHub | noreply@github.com | 2025-07-23T11:16:41+03:00 | | memory : handle saving/loading null layers in recurrent memory (#14675) |
| 1123 | 6c88b3bb2509d980e6a64c50fdf8dd304929f770 | 14c28dfc50a2e3915e697122cd3c5343224009be | lixing-star | 104126818+lixing-star@users.noreply.github.com | 2025-07-23T14:39:51+08:00 | GitHub | noreply@github.com | 2025-07-23T09:39:51+03:00 | | ggml: fix loongarch quantize_row_q8_1 error (#14827) |
| 1124 | 14c28dfc50a2e3915e697122cd3c5343224009be | 8c988fa41db4d9dbc05abfd666c04def1259c645 | chen fan | 350211548@qq.com | 2025-07-23T11:58:00+08:00 | GitHub | noreply@github.com | 2025-07-23T11:58:00+08:00 | | CANN: weight format to NZ for Ascend310P3 (#14407) |
| 1125 | 8c988fa41db4d9dbc05abfd666c04def1259c645 | acd6cb1c41676f6bbb25c2a76fa5abeb1719301e | Aman Gupta | amangupta052@gmail.com | 2025-07-23T09:25:42+08:00 | GitHub | noreply@github.com | 2025-07-23T09:25:42+08:00 | | CUDA: add fused rms norm (#14800) |
| 1126 | acd6cb1c41676f6bbb25c2a76fa5abeb1719301e | 84712b60439453fb393c2ca753cee682a4ad41f5 | Csaba Kecskemeti | csaba.kecskemeti@gmail.com | 2025-07-22T09:29:43-07:00 | GitHub | noreply@github.com | 2025-07-22T19:29:43+03:00 | | ggml : model card yaml tab->2xspace (#14819) |
| 1127 | 84712b60439453fb393c2ca753cee682a4ad41f5 | d4d1522b20809a350ffb094db20f40f17d3ab80f | Jeff Bolz | jbolz@nvidia.com | 2025-07-22T10:35:21-05:00 | GitHub | noreply@github.com | 2025-07-22T17:35:21+02:00 | | vulkan: fix rms_norm_mul to handle broadcasting dim0 (#14817) |
| 1128 | d4d1522b20809a350ffb094db20f40f17d3ab80f | d1aa0cc5d13ec3fa553de5d298f9166679e58479 | Molly Sophia | mollysophia379@gmail.com | 2025-07-22T23:01:29+08:00 | GitHub | noreply@github.com | 2025-07-22T23:01:29+08:00 | | llama : add model type detection for rwkv7 7B&14B (#14816) |
| 1129 | d1aa0cc5d13ec3fa553de5d298f9166679e58479 | c8ade30036139e32108fee53d8b7164dbfda4bee | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-22T13:33:37+01:00 | GitHub | noreply@github.com | 2025-07-22T14:33:37+02:00 | | imatrix: add option to display importance score statistics for a given imatrix file (#12718) |
| 1130 | c8ade30036139e32108fee53d8b7164dbfda4bee | e28c0b80c249b5618c3a0d5e960f63929f380ac9 | stduhpf | stephduh@live.fr | 2025-07-22T12:51:03+02:00 | GitHub | noreply@github.com | 2025-07-22T12:51:03+02:00 | | Mtmd: add a way to select device for vision encoder (#14236) |
| 1131 | e28c0b80c249b5618c3a0d5e960f63929f380ac9 | 8e6f8bc875358968b63e08f7bbbe0a288f29d856 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-22T12:33:10+02:00 | GitHub | noreply@github.com | 2025-07-22T12:33:10+02:00 | | cuda : implement bf16 cpy ops and enable bf16 cont (#14763) |
| 1132 | 8e6f8bc875358968b63e08f7bbbe0a288f29d856 | adef81781a15083f218eae6c488b95cdad781971 | lhez | lih@qti.qualcomm.com | 2025-07-21T23:53:30-07:00 | GitHub | noreply@github.com | 2025-07-22T08:53:30+02:00 | | opencl: remove unreachable `return` (#14806) |
| 1133 | adef81781a15083f218eae6c488b95cdad781971 | 48b86c4fdb1b1246d9fe17d2d53e3a1c2a0c3245 | Molly Sophia | mollysophia379@gmail.com | 2025-07-22T09:24:22+08:00 | GitHub | noreply@github.com | 2025-07-22T09:24:22+08:00 | | server : allow setting `--reverse-prompt` arg (#14799) |
| 1134 | 48b86c4fdb1b1246d9fe17d2d53e3a1c2a0c3245 | 38d3af1b73302377111e88ddd725257638339286 | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-22T07:45:26+08:00 | GitHub | noreply@github.com | 2025-07-22T07:45:26+08:00 | | cuda: remove linking to cublasLt (#14790) |
| 1135 | 38d3af1b73302377111e88ddd725257638339286 | 6c9ee3b17e19dcc82ab93d52ae46fdd0226d4777 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-21T22:55:10+02:00 | GitHub | noreply@github.com | 2025-07-21T13:55:10-07:00 | | opencl: fix `im2col` when `KW!=KH` (#14803) |
| 1136 | 6c9ee3b17e19dcc82ab93d52ae46fdd0226d4777 | cd465d823c378853f1f6570eebfb77f69c7e1d39 | rmatif | kingrealriadh@gmail.com | 2025-07-21T19:03:19+02:00 | GitHub | noreply@github.com | 2025-07-21T10:03:19-07:00 | | opencl: add conv2d kernel (#14403) |
| 1137 | cd465d823c378853f1f6570eebfb77f69c7e1d39 | 922042601b8a16877ccb1c2afaa2071f76734f10 | Romain Biessy | romain.biessy@codeplay.com | 2025-07-21T18:39:29+02:00 | GitHub | noreply@github.com | 2025-07-21T18:39:29+02:00 | | sycl: Fix im2col (#14797) |
| 1138 | 922042601b8a16877ccb1c2afaa2071f76734f10 | 2ba1333b35e471b344974dde553db11bf1c2836f | Charles Xu | charles.xu@arm.com | 2025-07-21T15:49:52+02:00 | GitHub | noreply@github.com | 2025-07-21T16:49:52+03:00 | | kleidiai: add support for get_rows (#14676) |
| 1139 | 2ba1333b35e471b344974dde553db11bf1c2836f | c2e058f1b4e799f1be085560c1bcef95b7b5ed02 | Radoslav Gerganov | rgerganov@gmail.com | 2025-07-21T15:03:49+03:00 | GitHub | noreply@github.com | 2025-07-21T14:03:49+02:00 | | docs : fix backends table in README.md (#14796) |
| 1140 | c2e058f1b4e799f1be085560c1bcef95b7b5ed02 | c82d48ec23fb8749c341d0838f6891fd5f6b6da0 | Jeff Bolz | jbolz@nvidia.com | 2025-07-21T06:35:40-05:00 | GitHub | noreply@github.com | 2025-07-21T13:35:40+02:00 | | vulkan/cuda: Fix im2col when KW!=KH (#14789) |
| 1141 | c82d48ec23fb8749c341d0838f6891fd5f6b6da0 | b4efd77f8ab407836ca73a5176f041650c5b2411 | Molly Sophia | mollysophia379@gmail.com | 2025-07-21T17:38:36+08:00 | GitHub | noreply@github.com | 2025-07-21T17:38:36+08:00 | | llama : fix `--reverse-prompt` crashing issue (#14794) |
| 1142 | b4efd77f8ab407836ca73a5176f041650c5b2411 | 2be60cbc2707359241c2784f9d2e30d8fc7cdabb | IsaacDynamo | 61521674+IsaacDynamo@users.noreply.github.com | 2025-07-21T09:24:51+02:00 | GitHub | noreply@github.com | 2025-07-21T10:24:51+03:00 | | server : add parse_special option to /tokenize endpoint (#14783) |
| 1143 | 2be60cbc2707359241c2784f9d2e30d8fc7cdabb | b526ad2668944a7b2b1721f60679153646313831 | Aman Gupta | amangupta052@gmail.com | 2025-07-21T02:13:47+08:00 | GitHub | noreply@github.com | 2025-07-20T20:13:47+02:00 | | docs : fix link for tools/perplexity in README.md (#14780) |
| 1144 | b526ad2668944a7b2b1721f60679153646313831 | 938b785764683c298e7805e712f8728489cc2f18 | rspOverflow | 217881046+rspOverflow@users.noreply.github.com | 2025-07-20T23:55:32+07:00 | GitHub | noreply@github.com | 2025-07-20T18:55:32+02:00 | | Documentation: Further revisions to the Vulkan section in build.md (#14785) |
| 1145 | 938b785764683c298e7805e712f8728489cc2f18 | 36c153248faf969af1b62ab231348694b2047b8b | Aman Gupta | amangupta052@gmail.com | 2025-07-20T19:42:34+08:00 | GitHub | noreply@github.com | 2025-07-20T19:42:34+08:00 | | Clang-format: local files first + fix BinPacking (#14779) |
| 1146 | 36c153248faf969af1b62ab231348694b2047b8b | a979ca22db0d737af1e548a73291193655c6be99 | 0cc4m | picard12@live.de | 2025-07-19T22:47:21+02:00 | GitHub | noreply@github.com | 2025-07-19T23:47:21+03:00 | | Contrib: add 0cc4m as codeowner for Vulkan backend (#14775) |
| 1147 | a979ca22db0d737af1e548a73291193655c6be99 | 90083283ec254fa8d33897746dea229aee401b37 | Ervin Áron Tasnádi | etasnadi@protonmail.com | 2025-07-19T21:59:08+02:00 | GitHub | noreply@github.com | 2025-07-19T21:59:08+02:00 | | ggml: adds CONV_2D op and direct GEMM Vulkan implementation (#14316) |
| 1148 | 90083283ec254fa8d33897746dea229aee401b37 | d4b91ea7b2da253e1355b503f0fcb7b428ce005d | compilade | git@compilade.net | 2025-07-19T12:51:22-04:00 | GitHub | noreply@github.com | 2025-07-19T12:51:22-04:00 | | imatrix : use GGUF to store importance matrices (#9400) |
| 1149 | d4b91ea7b2da253e1355b503f0fcb7b428ce005d | 83f5872404baa39d826af2ef66351e63c64205a8 | Peter0x44 | peter0x44@disroot.org | 2025-07-19T16:58:03+01:00 | GitHub | noreply@github.com | 2025-07-19T17:58:03+02:00 | | vulkan: Add logging for bf16 features to ggml_vk_print_gpu_info (#13274) (#14707) |
| 1150 | 83f5872404baa39d826af2ef66351e63c64205a8 | f0d4d176df72734a543c29eef9f942850c13311e | 0cc4m | picard12@live.de | 2025-07-19T17:47:53+02:00 | GitHub | noreply@github.com | 2025-07-19T17:47:53+02:00 | | Vulkan: Fix fprintf format-security warning (#14770) |
| 1151 | f0d4d176df72734a543c29eef9f942850c13311e | b17230917c18a25af9cd143a941001466af845a2 | rspOverflow | 217881046+rspOverflow@users.noreply.github.com | 2025-07-19T17:18:36+07:00 | GitHub | noreply@github.com | 2025-07-19T12:18:36+02:00 | | Documentation: Update build.md's Vulkan section (#14736) |
| 1152 | b17230917c18a25af9cd143a941001466af845a2 | bf9087f59aab940cf312b85a67067ce33d9e365a | Georgi Gerganov | ggerganov@gmail.com | 2025-07-19T11:46:12+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-19T11:46:50+03:00 | | sync : ggml |
| 1153 | bf9087f59aab940cf312b85a67067ce33d9e365a | 9fb1042ce6719bc46dbe88ab013148aabe3105f1 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T20:37:26+03:00 | GitHub | noreply@github.com | 2025-07-18T20:37:26+03:00 | | metal : fuse add, mul + add tests (#14596) |
| 1154 | 9fb1042ce6719bc46dbe88ab013148aabe3105f1 | 2adf8d83acdb9b1bf58db6c9729ac9dc6847a58b | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T20:08:33+03:00 | GitHub | noreply@github.com | 2025-07-18T20:08:33+03:00 | | graph : fix graph reuse reset of params (#14760) |
| 1155 | 2adf8d83acdb9b1bf58db6c9729ac9dc6847a58b | 021cc28bef4dd7d0bf9c91dbbd0803caa6cb15f2 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T17:33:41+03:00 | GitHub | noreply@github.com | 2025-07-18T17:33:41+03:00 | | parallel : add option for different RNG seeds (#14757) |
| 1156 | 021cc28bef4dd7d0bf9c91dbbd0803caa6cb15f2 | d498af3d5a00f96bdd37b534860f03a6d9e98d39 | Oliver Simons | oliver.simons@posteo.de | 2025-07-18T13:35:32+02:00 | GitHub | noreply@github.com | 2025-07-18T04:35:32-07:00 | | cuda : Fix Gemma3n not executed as CUDA_GRAPH on NVGPUs (#14741) |
| 1157 | d498af3d5a00f96bdd37b534860f03a6d9e98d39 | eacdeb5bfcb6c6cd54461fd0e9f04cab78bf975b | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T14:31:15+03:00 | GitHub | noreply@github.com | 2025-07-18T14:31:15+03:00 | | graph : avoid huge warm-up graphs for MoE models (#14753) |
| 1158 | eacdeb5bfcb6c6cd54461fd0e9f04cab78bf975b | e0cb5c5cb8a61ac232130cf6bf878035f93824d9 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T11:53:55+03:00 | GitHub | noreply@github.com | 2025-07-18T11:53:55+03:00 | | model : fix build after merge conflict (#14754) |
| 1159 | e0cb5c5cb8a61ac232130cf6bf878035f93824d9 | f9a31eea06a859e34cecb88b4d020c7f03d86cc4 | lgai-exaone | exaonemodels@lgresearch.ai | 2025-07-18T17:45:49+09:00 | GitHub | noreply@github.com | 2025-07-18T10:45:49+02:00 | | model : add EXAONE 4.0 support (#14630) |
| 1160 | f9a31eea06a859e34cecb88b4d020c7f03d86cc4 | 8f974bc1e980c06833504276021072e7a4088c81 | Aman Gupta | amangupta052@gmail.com | 2025-07-18T14:54:18+08:00 | GitHub | noreply@github.com | 2025-07-18T14:54:18+08:00 | | CUDA: set_rows + cpy.cu refactor (#14712) |
| 1161 | 8f974bc1e980c06833504276021072e7a4088c81 | 09651d09ffc1e941bd1be23163abf5495c416547 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-18T08:29:28+03:00 | GitHub | noreply@github.com | 2025-07-18T08:29:28+03:00 | | graph : refactor context to not pass gf explicitly (#14629) |
| 1162 | 09651d09ffc1e941bd1be23163abf5495c416547 | 349ea79fcebc75b0c55bf61594a47736966d4f95 | Nexes the Elder | 124105151+Nexesenex@users.noreply.github.com | 2025-07-18T06:25:54+02:00 | GitHub | noreply@github.com | 2025-07-18T07:25:54+03:00 | | graph : Pass the graph placeholder message in debug mode (#14748) |
| 1163 | 349ea79fcebc75b0c55bf61594a47736966d4f95 | 670e1360cd40f242ae76ba0966542fae6cb59392 | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-07-18T10:23:14+08:00 | GitHub | noreply@github.com | 2025-07-18T10:23:14+08:00 | | use max work group size for device to replace the magic number (#14732) |
| 1164 | 670e1360cd40f242ae76ba0966542fae6cb59392 | 760b4484e3c192a2649a6ffd0d90086c5558f849 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-07-18T01:17:16+02:00 | GitHub | noreply@github.com | 2025-07-18T01:17:16+02:00 | | convert : fix Ernie4.5 MoE without shared experts (#14746) |
| 1165 | 760b4484e3c192a2649a6ffd0d90086c5558f849 | cb887f1bc1001c92f7b4a595b9014f3a454a07ab | Wroclaw | wroclaw223@outlook.com | 2025-07-18T00:18:16+02:00 | GitHub | noreply@github.com | 2025-07-17T15:18:16-07:00 | | nix : use optionalAttrs for env mkDerivation attrset argument (#14726) |
| 1166 | cb887f1bc1001c92f7b4a595b9014f3a454a07ab | d6fb3f6b49b27ef1c0f4cf5128e041f7e7dc03af | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-07-17T23:15:32+02:00 | GitHub | noreply@github.com | 2025-07-17T23:15:32+02:00 | | model: add Ernie 4.5 MoE support (#14658) |
| 1167 | d6fb3f6b49b27ef1c0f4cf5128e041f7e7dc03af | 01612b74090df592663cfa01f661c9628f403b59 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-17T20:52:33+03:00 | GitHub | noreply@github.com | 2025-07-17T20:52:33+03:00 | | kv-cache : fix k-shift for multiple streams (#14742) |
| 1168 | 01612b74090df592663cfa01f661c9628f403b59 | 086cf81e88fb75287b71ff19c08a206b7bc2e02f | Georgi Gerganov | ggerganov@gmail.com | 2025-07-17T19:08:33+03:00 | GitHub | noreply@github.com | 2025-07-17T19:08:33+03:00 | | llama : reuse compute graphs (#14482) |
| 1169 | 086cf81e88fb75287b71ff19c08a206b7bc2e02f | d9b691081c04ec5fb0daa9d2b979f915c142963d | Tarek Dakhran | tarek@liquid.ai | 2025-07-17T09:22:11+02:00 | GitHub | noreply@github.com | 2025-07-17T09:22:11+02:00 | | llama : fix parallel processing for lfm2 (#14705) |
| 1170 | d9b691081c04ec5fb0daa9d2b979f915c142963d | ad57d3edd2f48cf6dc41a98fd9b303435ecb4fb0 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-17T09:49:15+03:00 | GitHub | noreply@github.com | 2025-07-17T09:49:15+03:00 | | kv-cache : opt mask set input (#14600) |
| 1171 | ad57d3edd2f48cf6dc41a98fd9b303435ecb4fb0 | 1ba45d49822c39ca9a552c7b75efe0495ff400c3 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-17T09:45:54+03:00 | GitHub | noreply@github.com | 2025-07-17T09:45:54+03:00 | | batch : fix uninitialized has_cpl flag (#14733) |
| 1172 | 1ba45d49822c39ca9a552c7b75efe0495ff400c3 | 19e5943d9e976b59ccd5a0a87ae3d2ff3da44390 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-17T01:52:08+02:00 | GitHub | noreply@github.com | 2025-07-16T20:52:08-03:00 | | ci : disable failing vulkan crossbuilds (#14723) |
| 1173 | 19e5943d9e976b59ccd5a0a87ae3d2ff3da44390 | 496957e1cbcb522abc63aa18521036e40efce985 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-16T23:17:43+02:00 | GitHub | noreply@github.com | 2025-07-16T23:17:43+02:00 | | convert : make hf token optional (#14717) |
| 1174 | 496957e1cbcb522abc63aa18521036e40efce985 | 21c021745d781edf9c44b4972ef6cbbf53b0ecff | Diner Burger | burger@diner.name | 2025-07-16T15:17:25-04:00 | GitHub | noreply@github.com | 2025-07-16T21:17:25+02:00 | | llama : fix parameter order for hybrid memory initialization (#14725) |
| 1175 | 21c021745d781edf9c44b4972ef6cbbf53b0ecff | b0f0ecc3dce806c68609d375a2b3edc430d8db18 | Reese Levine | reeselevine1@gmail.com | 2025-07-16T08:18:51-07:00 | GitHub | noreply@github.com | 2025-07-16T18:18:51+03:00 | | ggml: Add initial WebGPU backend (#14521) |
| 1176 | b0f0ecc3dce806c68609d375a2b3edc430d8db18 | 225e7a1438f4ea85eaa7b5ef3ab3b266ee4d9c06 | tempstudio | 49735574+tempstudio@users.noreply.github.com | 2025-07-16T10:02:06-05:00 | GitHub | noreply@github.com | 2025-07-16T18:02:06+03:00 | | model : support output bias for qwen2 (#14711) |
| 1177 | 225e7a1438f4ea85eaa7b5ef3ab3b266ee4d9c06 | ab140198211385b85eeeb0abd549a4bbe259e10d | Georgi Gerganov | ggerganov@gmail.com | 2025-07-16T16:35:42+03:00 | GitHub | noreply@github.com | 2025-07-16T16:35:42+03:00 | | llama : add high-throughput mode (#14363) |
| 1178 | ab140198211385b85eeeb0abd549a4bbe259e10d | 64978340b0b4a0a6e2fb74270c1509383d2eff32 | Aman Gupta | amangupta052@gmail.com | 2025-07-16T20:03:51+08:00 | GitHub | noreply@github.com | 2025-07-16T20:03:51+08:00 | | Support diffusion models: Add Dream 7B (#14644) |
| 1179 | 64978340b0b4a0a6e2fb74270c1509383d2eff32 | 6ffd4e9c442e99afac3d138543ebf86d5fc5ee03 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-16T14:43:32+03:00 | GitHub | noreply@github.com | 2025-07-16T14:43:32+03:00 | | ggml : add asserts (#14720) |
| 1180 | 6ffd4e9c442e99afac3d138543ebf86d5fc5ee03 | e4841d24d3485ea4af54f8b65a19ec3123f0ff3c | Georgi Gerganov | ggerganov@gmail.com | 2025-07-16T14:04:12+03:00 | GitHub | noreply@github.com | 2025-07-16T14:04:12+03:00 | | server : pre-calculate EOG logit biases (#14721) |
| 1181 | e4841d24d3485ea4af54f8b65a19ec3123f0ff3c | 538cc77f7f44dfa047dba6a06d90c86dda69cf1d | Shunta Saito | shunta.saito@gmail.com | 2025-07-16T19:12:22+09:00 | GitHub | noreply@github.com | 2025-07-16T12:12:22+02:00 | | llama : fix parallel processing for plamo2 (#14716) |
| 1182 | 538cc77f7f44dfa047dba6a06d90c86dda69cf1d | 5cae76654113160f691f581930b69fc5535e8159 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-16T12:13:57+03:00 | GitHub | noreply@github.com | 2025-07-16T12:13:57+03:00 | | server : fix handling of the ignore_eos flag (#14710) |
| 1183 | 5cae76654113160f691f581930b69fc5535e8159 | 4b91d6f71f14040979bcdb7b6729b3bca93ec1c1 | Johannes Gäßler | johannesg@5d6.de | 2025-07-16T09:33:28+02:00 | GitHub | noreply@github.com | 2025-07-16T09:33:28+02:00 | | scripts: synthetic prompt mode for server-bench.py (#14695) |
| 1184 | 4b91d6f71f14040979bcdb7b6729b3bca93ec1c1 | cf91f217f1c8b98b6db8e4ba6b480f017f81d206 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-16T08:52:04+02:00 | GitHub | noreply@github.com | 2025-07-16T08:52:04+02:00 | | convert : only check for tokenizer folder if we need it (#14704) |
| 1185 | cf91f217f1c8b98b6db8e4ba6b480f017f81d206 | 79e0b68c178656bb0632cb8602d2940b755077f8 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-16T08:51:12+02:00 | GitHub | noreply@github.com | 2025-07-16T08:51:12+02:00 | | convert : add pre-computed hashes first to prevent order mishaps (#14701) |
| 1186 | 79e0b68c178656bb0632cb8602d2940b755077f8 | c81f4192f91a1e209c1eec7a84fe5371ef9175da | Min-Hua | 136287195+Min-Hua@users.noreply.github.com | 2025-07-16T12:00:42+08:00 | GitHub | noreply@github.com | 2025-07-16T07:00:42+03:00 | | llama: add LLAMA_API to deprecated llama_kv_self_seq_div (#14708) |
| 1187 | c81f4192f91a1e209c1eec7a84fe5371ef9175da | 4a4f426944e79b79e389f9ed7b34831cb9b637ad | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-15T23:04:42+01:00 | GitHub | noreply@github.com | 2025-07-16T00:04:42+02:00 | | gguf-py : dump bpw per layer and model in markdown mode (#14703) |
| 1188 | 4a4f426944e79b79e389f9ed7b34831cb9b637ad | ba1ceb34566c889a1fc500efa79799ffed25d9b0 | Gabriel Larson | 55459720+gabriellarson@users.noreply.github.com | 2025-07-15T14:54:22-05:00 | GitHub | noreply@github.com | 2025-07-15T21:54:22+02:00 | | model : add Kimi-K2 support (#14654) |
| 1189 | ba1ceb34566c889a1fc500efa79799ffed25d9b0 | 10a0351a97c25471aea0bbde9cca54d32d163eec | Jeff Bolz | jbolz@nvidia.com | 2025-07-15T14:51:09-05:00 | GitHub | noreply@github.com | 2025-07-15T21:51:09+02:00 | | vulkan: fix noncontig check for mat_mul_id splitting (#14683) |
| 1190 | 10a0351a97c25471aea0bbde9cca54d32d163eec | 68e37a61a7b24863f67541db65b4f6195962a268 | Jeff Bolz | jbolz@nvidia.com | 2025-07-15T14:32:11-05:00 | GitHub | noreply@github.com | 2025-07-15T21:32:11+02:00 | | vulkan: add RTE variants for glu/add/sub/mul/div (#14653) |
| 1191 | 68e37a61a7b24863f67541db65b4f6195962a268 | cbc68be51d88b1d5531643b926a4b359c3cff131 | Shunta Saito | shunta.saito@gmail.com | 2025-07-16T01:11:42+09:00 | GitHub | noreply@github.com | 2025-07-15T18:11:42+02:00 | | model : add PLaMo-2 support (#14560) |
| 1192 | cbc68be51d88b1d5531643b926a4b359c3cff131 | bdca38376f7e8dd928defe01ce6a16218a64b040 | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-15T15:28:53+08:00 | GitHub | noreply@github.com | 2025-07-15T15:28:53+08:00 | | cuda: fix build warnings in set-rows.cu (unused variable) (#14687) |
| 1193 | bdca38376f7e8dd928defe01ce6a16218a64b040 | 55c509daf51d25bfaee9c8b8ce6abff103d4473b | Anton Mitkov | anton_b_mitkov@abv.bg | 2025-07-14T18:12:42+01:00 | GitHub | noreply@github.com | 2025-07-14T18:12:42+01:00 | | sycl: Hotfix for non dnnl codepath (#14677) |
| 1194 | 55c509daf51d25bfaee9c8b8ce6abff103d4473b | 9c9e4fc6354fc811efa06a8eb7a86d3315cec9c8 | shalinib-ibm | Shalini.Salomi.Bodapati@ibm.com | 2025-07-14T18:46:42+05:30 | GitHub | noreply@github.com | 2025-07-14T16:16:42+03:00 | | ggml : refactor llamafile_sgemm PPC code (#14673) |
| 1195 | 9c9e4fc6354fc811efa06a8eb7a86d3315cec9c8 | 494c5899cb76859f32ddd913534f2685fd684a3d | Aman Gupta | amangupta052@gmail.com | 2025-07-14T21:01:41+08:00 | GitHub | noreply@github.com | 2025-07-14T21:01:41+08:00 | | llama-context: add ability to get logits (#14672) |
| 1196 | 494c5899cb76859f32ddd913534f2685fd684a3d | 0f4c6ec0f1a9607ba67071f8a02c69b0afc2f91e | Johannes Gäßler | johannesg@5d6.de | 2025-07-14T13:14:30+02:00 | GitHub | noreply@github.com | 2025-07-14T13:14:30+02:00 | | scripts: benchmark for HTTP server throughput (#14668) |
| 1197 | 0f4c6ec0f1a9607ba67071f8a02c69b0afc2f91e | 65a3ebb0aa56d6c501466e0f950ad15105fd32d8 | Akarshan Biswas | akarshan@menlo.ai | 2025-07-14T15:07:55+05:30 | GitHub | noreply@github.com | 2025-07-14T10:37:55+01:00 | | SYCL: use 1D kernel for set_rows (#14618) |
| 1198 | 65a3ebb0aa56d6c501466e0f950ad15105fd32d8 | 0d9226763c82562186122f3b827fa3862864a19c | Anton Mitkov | anton_b_mitkov@abv.bg | 2025-07-14T10:37:35+01:00 | GitHub | noreply@github.com | 2025-07-14T10:37:35+01:00 | | sycl: Batched mulmat rework for oneDNN dispatch (#14617) |
| 1199 | 0d9226763c82562186122f3b827fa3862864a19c | 982e347255723fe6d02e60ee30cfdd0559c884c5 | Molly Sophia | mollysophia379@gmail.com | 2025-07-14T07:43:43+08:00 | GitHub | noreply@github.com | 2025-07-14T07:43:43+08:00 | | llama : add jinja template for rwkv-world (#14665) |
| 1200 | 982e347255723fe6d02e60ee30cfdd0559c884c5 | 923e3ea2e3c96a0b4c208f53bc3bc90cdcdf13c0 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-07-13T17:02:17+01:00 | GitHub | noreply@github.com | 2025-07-13T18:02:17+02:00 | | quantize : fix minor logic flaw in --tensor-type (#14572) |
| 1201 | 923e3ea2e3c96a0b4c208f53bc3bc90cdcdf13c0 | e743cddb60dc3a8815b9de7dd7d5c491e61b2259 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-13T15:01:24+02:00 | GitHub | noreply@github.com | 2025-07-13T15:01:24+02:00 | | cuda : add set rows for bf16 (#14664) |
| 1202 | e743cddb60dc3a8815b9de7dd7d5c491e61b2259 | 05fec5bd298d3c0243cbb9336e59b8b6aff75a81 | Yavor Ivanov | yavorgenadiev@gmail.com | 2025-07-13T02:33:16-07:00 | GitHub | noreply@github.com | 2025-07-13T11:33:16+02:00 | | cuda : add ELU support (#14657) |
| 1203 | 05fec5bd298d3c0243cbb9336e59b8b6aff75a81 | dcf7f2ea3c4e3cf36dd9ab5a36785c00e6033267 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-13T10:36:33+03:00 | GitHub | noreply@github.com | 2025-07-13T10:36:33+03:00 | | ggml : add build-time message to remind about ggml_set_rows (#14661) |
| 1204 | dcf7f2ea3c4e3cf36dd9ab5a36785c00e6033267 | 84b396e0510855a95d591afdf1f21c562cb3712a | Yavor Ivanov | yavorgenadiev@gmail.com | 2025-07-12T22:38:13-07:00 | GitHub | noreply@github.com | 2025-07-13T08:38:13+03:00 | | metal : Add missing unary ops Metal support (#14660) |
| 1205 | 84b396e0510855a95d591afdf1f21c562cb3712a | c31e60647def83d671bac5ab5b35579bf25d9aa1 | Yavor Ivanov | yavorgenadiev@gmail.com | 2025-07-12T22:12:36-07:00 | GitHub | noreply@github.com | 2025-07-13T08:12:36+03:00 | | cmake : Add CMake presets for Linux and GCC (#14656) |
| 1206 | c31e60647def83d671bac5ab5b35579bf25d9aa1 | 67eade1bf93f40e7c4975e5bd0782f426c89cff4 | Tarek Dakhran | tarek@liquid.ai | 2025-07-12T19:10:14+02:00 | GitHub | noreply@github.com | 2025-07-12T19:10:14+02:00 | | tests : cover lfm2 cases in test_ssm_conv (#14651) |
| 1207 | 67eade1bf93f40e7c4975e5bd0782f426c89cff4 | 7de5c7cab61d4da4387ed9b216f88b96297bcc2d | Tarek Dakhran | tarek@liquid.ai | 2025-07-12T19:07:08+02:00 | GitHub | noreply@github.com | 2025-07-12T19:07:08+02:00 | | docs : add LFM2 to models section (#14650) |
| 1208 | 7de5c7cab61d4da4387ed9b216f88b96297bcc2d | 8eff95544e817704d44bec20f9fc956ce76a33be | Aman Gupta | amangupta052@gmail.com | 2025-07-12T21:31:38+08:00 | GitHub | noreply@github.com | 2025-07-12T16:31:38+03:00 | | CUDA: add set rows for f32 and f16 (#14551) |
| 1209 | 8eff95544e817704d44bec20f9fc956ce76a33be | 3120413ccd7797d4cbdee32ef89a641765d1f6c4 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T16:06:12+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T16:13:27+03:00 | | sync : ggml |
| 1210 | 3120413ccd7797d4cbdee32ef89a641765d1f6c4 | 215535701d659e873c52d2a9163d4616a42da4f7 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T12:39:32+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T14:25:44+03:00 | | vulkan : remove unused vars (#0) |
| 1211 | 215535701d659e873c52d2a9163d4616a42da4f7 | 74bb294591054a6bf0cc3f6e4637d87414cf2c46 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T12:39:27+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T14:25:44+03:00 | | sync : ggml |
| 1212 | 74bb294591054a6bf0cc3f6e4637d87414cf2c46 | 3e303b1107e4748cd50a675544d8c7fabb40ec54 | Acly | aclysia@gmail.com | 2025-07-12T12:37:37+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T14:25:44+03:00 | | vulkan : implement bilinear interpolation (ggml/1291) |
| 1213 | 3e303b1107e4748cd50a675544d8c7fabb40ec54 | 0c1df14b5f8d992805cb22d0b77b44092a18aeab | Acly | aclysia@gmail.com | 2025-07-12T12:32:32+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-12T14:25:44+03:00 | | vulkan : implement ggml_roll (ggml/1290) |
| 1214 | 0c1df14b5f8d992805cb22d0b77b44092a18aeab | b3ad3a0191994d6c47b2bd389d5c7431526ecd2c | Douglas Hanley | thesecretaryofwar@gmail.com | 2025-07-12T06:21:02-04:00 | GitHub | noreply@github.com | 2025-07-12T13:21:02+03:00 | | server : fix pooled embedding output (#14645) |
| 1215 | b3ad3a0191994d6c47b2bd389d5c7431526ecd2c | 98197e5c98388470030d908f355ec5937dcccaaa | Jeff Bolz | jbolz@nvidia.com | 2025-07-12T05:12:26-05:00 | GitHub | noreply@github.com | 2025-07-12T12:12:26+02:00 | | vulkan: support SET_ROWS (#14587) |
| 1216 | 98197e5c98388470030d908f355ec5937dcccaaa | f5e96b368f1acc7f53c390001b936517c4d18999 | Jeff Bolz | jbolz@nvidia.com | 2025-07-12T04:51:58-05:00 | GitHub | noreply@github.com | 2025-07-12T11:51:58+02:00 | | vulkan: optimizations for deepseek prompt processing (#14555) |
| 1217 | f5e96b368f1acc7f53c390001b936517c4d18999 | 756aa1020aabc70b71b9a8efcd35fbf42b3bb9ab | Tarek Dakhran | t.dakhran@gmail.com | 2025-07-11T20:27:01+02:00 | GitHub | noreply@github.com | 2025-07-11T20:27:01+02:00 | | model : support LiquidAI LFM2 hybrid family (#14620) |
| 1218 | 756aa1020aabc70b71b9a8efcd35fbf42b3bb9ab | aaa088d87f9006d56866085fe46e4b2755ef723f | Slobodan Josic | 127323561+slojosic-amd@users.noreply.github.com | 2025-07-11T18:55:00+02:00 | GitHub | noreply@github.com | 2025-07-11T18:55:00+02:00 | | HIP : Add HIP 7.0+ compatibility for hipBLAS compute types (#14634) |
| 1219 | aaa088d87f9006d56866085fe46e4b2755ef723f | 0d5375d54b258ec63edd1fb5d58c37d58ce8be8b | Georgi Gerganov | ggerganov@gmail.com | 2025-07-11T16:07:55+03:00 | GitHub | noreply@github.com | 2025-07-11T16:07:55+03:00 | | readme : add hot PRs (#14636) |
| 1220 | 0d5375d54b258ec63edd1fb5d58c37d58ce8be8b | 576c82eda210ca0111c04f5256bf77897a4d4cc4 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-11T13:46:07+03:00 | GitHub | noreply@github.com | 2025-07-11T13:46:07+03:00 | | llama : move enum llama_vocab_pre_type to implementation (#14631) |
| 1221 | 576c82eda210ca0111c04f5256bf77897a4d4cc4 | 0aedae00e6fb48680324a5ac5da9cba0e35de6b5 | Dowon | ks2515@naver.com | 2025-07-11T16:36:04+09:00 | GitHub | noreply@github.com | 2025-07-11T09:36:04+02:00 | | vocab : add midm-2.0 model pre-tokenizer (#14626) |
| 1222 | 0aedae00e6fb48680324a5ac5da9cba0e35de6b5 | 6bdda13981d6c8189b7dc4f9fb8ecb91c21529f8 | Gabe Goodhart | ghart@us.ibm.com | 2025-07-10T18:20:13-06:00 | GitHub | noreply@github.com | 2025-07-11T02:20:13+02:00 | | model : Granite Four (#13550) |
| 1223 | 6bdda13981d6c8189b7dc4f9fb8ecb91c21529f8 | 0b8855775c6b873931d40b77a5e42558aacbde52 | rmatif | kingrealriadh@gmail.com | 2025-07-10T23:58:12+02:00 | GitHub | noreply@github.com | 2025-07-10T14:58:12-07:00 | | opencl: add tiled mul_mat_f16_f32 (#14535) |
| 1224 | 0b8855775c6b873931d40b77a5e42558aacbde52 | 4bb625b713fd9b294b4f7af87eaa752b710c7cc1 | lhez | quic_lih@quicinc.com | 2025-07-10T11:48:52-07:00 | GitHub | noreply@github.com | 2025-07-10T11:48:52-07:00 | | opencl: add `set_rows` for `f16` and `f32` (#14547) |
| 1225 | 4bb625b713fd9b294b4f7af87eaa752b710c7cc1 | 11ee0fea2a24da1d3206eeba8fc52b759d9dfb24 | Ryan Mangeno | 160974989+ryan-mangeno@users.noreply.github.com | 2025-07-10T13:41:00-04:00 | GitHub | noreply@github.com | 2025-07-10T19:41:00+02:00 | | Smoldocling support (#14597) |
| 1226 | 11ee0fea2a24da1d3206eeba8fc52b759d9dfb24 | a457551332853ef19d0796fec12b62c538126ea5 | Aman Gupta | amangupta052@gmail.com | 2025-07-10T23:29:01+08:00 | GitHub | noreply@github.com | 2025-07-10T23:29:01+08:00 | | Docs: script to auto-generate ggml operations docs (#14598) |
| 1227 | a457551332853ef19d0796fec12b62c538126ea5 | 704bb7a71c01dc07c1478b85f6322bf5dfde1eaf | Eric Zhang | 34133756+EZForever@users.noreply.github.com | 2025-07-10T20:29:05+08:00 | GitHub | noreply@github.com | 2025-07-10T15:29:05+03:00 | | cmake : do not search for curl libraries by ourselves (#14613) |
| 1228 | 704bb7a71c01dc07c1478b85f6322bf5dfde1eaf | 435a6d10d618a015060e45a38c7e9f27f4243316 | Akarshan Biswas | akarshan@menlo.ai | 2025-07-10T13:59:38+05:30 | GitHub | noreply@github.com | 2025-07-10T09:29:38+01:00 | | SYCL: Initial set_rows kernel implementation (#14562) |
| 1229 | 435a6d10d618a015060e45a38c7e9f27f4243316 | f9a867f5921a85f3fa64d7b067f4c8ffc5f62eb4 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-10T09:00:20+02:00 | GitHub | noreply@github.com | 2025-07-10T10:00:20+03:00 | | llama : minor coding style fix for smollm3 (#14605) |
| 1230 | f9a867f5921a85f3fa64d7b067f4c8ffc5f62eb4 | ac44eb6c808bd5d677261ce86edd8c43ec54cf2c | Eric Zhang | 34133756+EZForever@users.noreply.github.com | 2025-07-10T13:19:37+08:00 | GitHub | noreply@github.com | 2025-07-10T08:19:37+03:00 | | cmake : bump llguidance version to v1.0.1 (#14609) |
| 1231 | ac44eb6c808bd5d677261ce86edd8c43ec54cf2c | a57d1bcb3c0165ac87b1f0dbb429839b0da69689 | Eric Zhang | 34133756+EZForever@users.noreply.github.com | 2025-07-10T13:19:13+08:00 | GitHub | noreply@github.com | 2025-07-10T08:19:13+03:00 | | cmake : llguidance build parser library only (#14608) |
| 1232 | a57d1bcb3c0165ac87b1f0dbb429839b0da69689 | cb9178f885d1986cc0b12feb26ff426bc8a3556c | compilade | git@compilade.net | 2025-07-09T23:54:38-04:00 | GitHub | noreply@github.com | 2025-07-09T23:54:38-04:00 | | cuda : support Falcon-H1 state size for SSM_SCAN (#14602) |
| 1233 | cb9178f885d1986cc0b12feb26ff426bc8a3556c | 4a5686da22057867c23bd4a6be941ddc8c51e585 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-09T23:09:28+02:00 | GitHub | noreply@github.com | 2025-07-09T23:09:28+02:00 | | llama : remove llm_graph_input_one (#14603) |
| 1234 | 4a5686da22057867c23bd4a6be941ddc8c51e585 | 98bab638fb28cf95a5a66dd2d51b40d6c8f6d69a | compilade | git@compilade.net | 2025-07-09T14:59:57-04:00 | GitHub | noreply@github.com | 2025-07-09T14:59:57-04:00 | | llama : support Jamba hybrid Transformer-Mamba models (#7531) |
| 1235 | 98bab638fb28cf95a5a66dd2d51b40d6c8f6d69a | 26a48ad699d50b6268900062661bd22f3e792579 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-09T18:16:12+02:00 | GitHub | noreply@github.com | 2025-07-09T18:16:12+02:00 | | ggml : add ggml_scale_bias (#14417) |
| 1236 | 26a48ad699d50b6268900062661bd22f3e792579 | ffd59e7d18a76459d5c31ba97073c7c9d73cb752 | Miaoqian Lin | linmq006@gmail.com | 2025-07-09T20:33:53+08:00 | GitHub | noreply@github.com | 2025-07-09T14:33:53+02:00 | | ggml : prevent integer overflow in gguf tensor size calculation (#14595) |
| 1237 | ffd59e7d18a76459d5c31ba97073c7c9d73cb752 | 105554595f9a7bf3e02232ed7798201d47c2a4a2 | Dowon | ks2515@naver.com | 2025-07-09T17:22:31+09:00 | GitHub | noreply@github.com | 2025-07-09T11:22:31+03:00 | | model : add skt/A.X-4.0 model vocabulary (#14589) |
| 1238 | 105554595f9a7bf3e02232ed7798201d47c2a4a2 | 04655063c47af6cbced295c8c7ad369402b15300 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-09T10:19:50+02:00 | GitHub | noreply@github.com | 2025-07-09T10:19:50+02:00 | | llama : remove unintended whitespace (#14592) |
| 1239 | 04655063c47af6cbced295c8c7ad369402b15300 | 20b7bf8a32259ad9189a4797fd3e3a859c537b99 | ibrahim khadraoui | 132432132+ibrahimkhadraoui@users.noreply.github.com | 2025-07-09T12:03:49+04:00 | GitHub | noreply@github.com | 2025-07-09T10:03:49+02:00 | | model : add support for Falcon-H1 family (#14534) |
| 1240 | 20b7bf8a32259ad9189a4797fd3e3a859c537b99 | 6efcd65945a98cf6883cdd9de4c8ccd8c79d219a | Xuan-Son Nguyen | son@huggingface.co | 2025-07-09T08:26:13+02:00 | GitHub | noreply@github.com | 2025-07-09T09:26:13+03:00 | | convert : fix smollm3 jinja template (#14586) |
| 1241 | 6efcd65945a98cf6883cdd9de4c8ccd8c79d219a | 699f4392a33f57c3352cf8d60bdc53db7ca235e7 | Jeff Bolz | jbolz@nvidia.com | 2025-07-08T13:11:42-05:00 | GitHub | noreply@github.com | 2025-07-08T20:11:42+02:00 | | vulkan: optimize flash attention split_k_reduce (#14554) |
| 1242 | 699f4392a33f57c3352cf8d60bdc53db7ca235e7 | 08382869a2d6dca5d84710f2f82cc17a9696585a | stevenkuang | stevenkuang@tencent.com | 2025-07-09T00:29:29+08:00 | GitHub | noreply@github.com | 2025-07-08T18:29:29+02:00 | | model : fix hunyuan moe chat template (#14584) |
| 1243 | 08382869a2d6dca5d84710f2f82cc17a9696585a | bb4f7a9e4eec171fecf0f640b1337a1c24485560 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-08T18:07:01+02:00 | GitHub | noreply@github.com | 2025-07-08T18:07:01+02:00 | | model : add SmolLM3 (#14581) |
| 1244 | bb4f7a9e4eec171fecf0f640b1337a1c24485560 | b8eeb8741d4483daf576498cf90537b4de71633c | compilade | git@compilade.net | 2025-07-08T11:37:47-04:00 | GitHub | noreply@github.com | 2025-07-08T18:37:47+03:00 | | memory : fix broken batch splits for recurrent cache (#14575) |
| 1245 | b8eeb8741d4483daf576498cf90537b4de71633c | 17a1f0d2d407040ee242e18dd79be8bb212cfcef | Jeff Bolz | jbolz@nvidia.com | 2025-07-08T08:21:21-05:00 | GitHub | noreply@github.com | 2025-07-08T15:21:21+02:00 | | vulkan : fix rope with partial rotation and non-cont src (#14582) |
| 1246 | 17a1f0d2d407040ee242e18dd79be8bb212cfcef | 8f22dc0a53338c629c1ef8fa878d8e39bfe627c9 | Alawode Oluwandabira | dabiraalawode@yahoo.com | 2025-07-08T11:47:33+03:00 | GitHub | noreply@github.com | 2025-07-08T11:47:33+03:00 | | server: Add ability to mount server at prefix (#14544) |
| 1247 | 8f22dc0a53338c629c1ef8fa878d8e39bfe627c9 | 53903ae6fa5f1caf187889c839cdd1ad25da4018 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-08T10:24:06+02:00 | GitHub | noreply@github.com | 2025-07-08T11:24:06+03:00 | | model : add hunyuan moe (#14425) |
| 1248 | 53903ae6fa5f1caf187889c839cdd1ad25da4018 | 4d0dcd4a06080e796e6742a88f2ffa7fc41b28b8 | Jeff Bolz | jbolz@nvidia.com | 2025-07-08T02:38:31-05:00 | GitHub | noreply@github.com | 2025-07-08T09:38:31+02:00 | | vulkan: increase timeout for CI (#14574) |
| 1249 | 4d0dcd4a06080e796e6742a88f2ffa7fc41b28b8 | 75c91de6e955d5b8f3f28173f5040593e1964eb3 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-08T10:15:21+03:00 | GitHub | noreply@github.com | 2025-07-08T10:15:21+03:00 | | cuda : fix rope with partial rotation and non-cont src (#14580) |
| 1250 | 75c91de6e955d5b8f3f28173f5040593e1964eb3 | 68155c66f0e76680f34442247e589a090add22d3 | Aman Gupta | amangupta052@gmail.com | 2025-07-08T10:11:18+08:00 | GitHub | noreply@github.com | 2025-07-08T10:11:18+08:00 | | CUDA: add bilinear interpolation for upscale (#14563) |
| 1251 | 68155c66f0e76680f34442247e589a090add22d3 | e1a7059053a9b8958f2d57a21fc46dbc7fb24f8e | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-08T07:58:30+08:00 | GitHub | noreply@github.com | 2025-07-08T07:58:30+08:00 | | musa: fix build warnings (unused variable) (#14561) |
| 1252 | e1a7059053a9b8958f2d57a21fc46dbc7fb24f8e | 12f55c302b35cfe900b84c5fe67c262026af9c44 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-07T23:35:35+02:00 | GitHub | noreply@github.com | 2025-07-07T23:35:35+02:00 | | llama : fix incorrect minicpm3 v_states shape (#14571) |
| 1253 | 12f55c302b35cfe900b84c5fe67c262026af9c44 | b9c3eefde1b67104bd993485ff38dd62abe9d70c | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-07T21:35:08+02:00 | GitHub | noreply@github.com | 2025-07-07T21:35:08+02:00 | | llama : remove ggml_cont where possible (#14568) |
| 1254 | b9c3eefde1b67104bd993485ff38dd62abe9d70c | 6491d6e4f1caf0ad2221865b4249ae6938a6308c | Aman Gupta | amangupta052@gmail.com | 2025-07-07T21:45:43+08:00 | GitHub | noreply@github.com | 2025-07-07T21:45:43+08:00 | | CUDA: add bf16 and i32 to getrows (#14529) |
| 1255 | 6491d6e4f1caf0ad2221865b4249ae6938a6308c | e592be15756ae546ba87bb5a5250b71248121971 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-07-06T10:29:36Z | GitHub | noreply@github.com | 2025-07-06T12:29:36+02:00 | | vulkan: increase LOAD_VEC_A to 8 (IQ1/IQ2) or 4 (IQ3) (#14485) |
| 1256 | e592be15756ae546ba87bb5a5250b71248121971 | a0374a67e2924f2e845cdc59dd67d9a44065a89c | Jeff Bolz | jbolz@nvidia.com | 2025-07-06T03:08:16-05:00 | GitHub | noreply@github.com | 2025-07-06T10:08:16+02:00 | | vulkan: fix rms_norm+mul fusion (#14545) |
| 1257 | a0374a67e2924f2e845cdc59dd67d9a44065a89c | ddef99522d1ba74193b7394e803fab8db5c78bae | Jeff Bolz | jbolz@nvidia.com | 2025-07-05T02:26:04-05:00 | GitHub | noreply@github.com | 2025-07-05T09:26:04+02:00 | | vulkan: Handle updated FA dim2/3 definition (#14518) |
| 1258 | ddef99522d1ba74193b7394e803fab8db5c78bae | 668168814601d2aef6161ce49ab186bc74177ae7 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-05T09:17:14+02:00 | GitHub | noreply@github.com | 2025-07-05T09:17:14+02:00 | | server : fix assistant prefilling when content is an array (#14360) |
| 1259 | 668168814601d2aef6161ce49ab186bc74177ae7 | bac8bed248d15419137c5bc7f834582397baaebc | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-05T08:24:56+02:00 | GitHub | noreply@github.com | 2025-07-04T23:24:56-07:00 | | opencl: add GELU_ERF (#14476) |
| 1260 | bac8bed248d15419137c5bc7f834582397baaebc | b81510a7b78f7c5a6f069c4f3d0e569b9d287913 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-05T07:18:09+03:00 | GitHub | noreply@github.com | 2025-07-05T07:18:09+03:00 | | eval-callback : check for empty input (#14539) |
| 1261 | b81510a7b78f7c5a6f069c4f3d0e569b9d287913 | ef797db357e44ecb7437fa9d22f4e1614104b342 | R0CKSTAR | yeahdongcn@gmail.com | 2025-07-05T12:10:53+08:00 | GitHub | noreply@github.com | 2025-07-05T12:10:53+08:00 | | test-backend-ops: add support for specifying output format (#14368) |
| 1262 | ef797db357e44ecb7437fa9d22f4e1614104b342 | 67d1ef23c68e1e09dc6d7423d1ef5fc0ea56ed5e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-04T19:19:09+03:00 | GitHub | noreply@github.com | 2025-07-04T19:19:09+03:00 | | metal : disable fast math in all quantize kernels (#14528) |
| 1263 | 67d1ef23c68e1e09dc6d7423d1ef5fc0ea56ed5e | 7b50f7c0257695d0f6918bffce4e68d5c13b0c9e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-04T09:08:59+03:00 | GitHub | noreply@github.com | 2025-07-04T09:08:59+03:00 | | batch : add optional for sequential equal split (#14511) |
| 1264 | 7b50f7c0257695d0f6918bffce4e68d5c13b0c9e | c79184d2d192489e3c918bab8ed717d22f8c02bd | Georgi Gerganov | ggerganov@gmail.com | 2025-07-04T09:05:36+03:00 | GitHub | noreply@github.com | 2025-07-04T09:05:36+03:00 | | graph : prepare for 4D mask (#14515) |
| 1265 | c79184d2d192489e3c918bab8ed717d22f8c02bd | 499a8f5a787f5fdc7f3fdfb9d1ff01b6cb5e994e | Georgi Gerganov | ggerganov@gmail.com | 2025-07-04T09:04:59+03:00 | GitHub | noreply@github.com | 2025-07-04T09:04:59+03:00 | | batch : add n_used count (#14512) |
| 1266 | 499a8f5a787f5fdc7f3fdfb9d1ff01b6cb5e994e | 28657a8229b5adc6028cf1c4ed62191792d2fdb0 | luyhcsu | 110711054+luyhcsu@users.noreply.github.com | 2025-07-04T11:50:07+08:00 | GitHub | noreply@github.com | 2025-07-04T11:50:07+08:00 | | CANN: Replace aclrtMemsetSync with aclnnInplaceZero operator (#14002) |
| 1267 | 28657a8229b5adc6028cf1c4ed62191792d2fdb0 | bee28421be25fd447f61cb6db64d556cbfce32ec | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-03T23:07:22+02:00 | GitHub | noreply@github.com | 2025-07-03T23:07:22+02:00 | | ggml : implement GEGLU_ERF and GEGLU_QUICK ops (#14445) |
| 1268 | bee28421be25fd447f61cb6db64d556cbfce32ec | 2b72bedec198a90bb5b0cceaf1d0aff9e34ffbc2 | lhez | quic_lih@quicinc.com | 2025-07-03T11:22:24-07:00 | GitHub | noreply@github.com | 2025-07-03T20:22:24+02:00 | | opencl : broadcast for soft_max (#14510) |
| 1269 | 2b72bedec198a90bb5b0cceaf1d0aff9e34ffbc2 | c8c4495b8d3a8799e2d46778f993965b0ac1ae43 | Jeff Bolz | jbolz@nvidia.com | 2025-07-03T13:21:14-05:00 | GitHub | noreply@github.com | 2025-07-03T20:21:14+02:00 | | vulkan: support mixed/deepseekR1 FA head sizes (#14509) |
| 1270 | c8c4495b8d3a8799e2d46778f993965b0ac1ae43 | 7b63a71a6b0f54effe9b94073d4d0519dcf53676 | Johannes Gäßler | johannesg@5d6.de | 2025-07-03T17:05:18+02:00 | GitHub | noreply@github.com | 2025-07-03T17:05:18+02:00 | | ggml: backward pass for split swiglu (#14483) |
| 1271 | 7b63a71a6b0f54effe9b94073d4d0519dcf53676 | 0c2ee38ab729f3f6a7c28dc4c5dcd8d5f0cbf8fc | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-07-03T11:00:03+02:00 | GitHub | noreply@github.com | 2025-07-03T11:00:03+02:00 | | Fix conditional enabling following arch checks for ggml-sycl (#14504) |
| 1272 | 0c2ee38ab729f3f6a7c28dc4c5dcd8d5f0cbf8fc | a70c8a0c4b4c1606cd9a0ba889ce61aa88610095 | Xuan-Son Nguyen | son@huggingface.co | 2025-07-03T10:03:06+02:00 | GitHub | noreply@github.com | 2025-07-03T10:03:06+02:00 | | convert : correct gemma 3n conversion (#14450) |
| 1273 | a70c8a0c4b4c1606cd9a0ba889ce61aa88610095 | 9067487c4411efb20400103fcccfdd389c80d428 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-03T10:53:35+03:00 | GitHub | noreply@github.com | 2025-07-03T10:53:35+03:00 | | kv-cache : use ggml_set_rows (#14285) |
| 1274 | 9067487c4411efb20400103fcccfdd389c80d428 | d4cdd9c1c3cebdafca735958597de4ff7b7c0f54 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-03T10:46:57+03:00 | GitHub | noreply@github.com | 2025-07-03T10:46:57+03:00 | | ggml : fix FA mask dim 2 and 3 (#14505) |
| 1275 | d4cdd9c1c3cebdafca735958597de4ff7b7c0f54 | 55c2646b452419be7552157b9b11373ccc1bf2f2 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-03T07:48:32+03:00 | GitHub | noreply@github.com | 2025-07-03T07:48:32+03:00 | | ggml : remove kompute backend (#14501) |
| 1276 | 55c2646b452419be7552157b9b11373ccc1bf2f2 | e75ba4c0434eb759eb7ff74e034ebe729053e575 | Aman Gupta | amangupta052@gmail.com | 2025-07-03T07:45:11+08:00 | GitHub | noreply@github.com | 2025-07-03T07:45:11+08:00 | | CUDA: add dynamic shared mem to softmax, refactor general usage (#14497) |
| 1277 | e75ba4c0434eb759eb7ff74e034ebe729053e575 | 5d46babdc2d4675d96ebcf23cac098a02f0d30cc | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-02T21:02:35+02:00 | GitHub | noreply@github.com | 2025-07-02T21:02:35+02:00 | | gguf-py : add support for chat template jinja files (#14508) |
| 1278 | 5d46babdc2d4675d96ebcf23cac098a02f0d30cc | e17991c466ac835b2c71bd813c1ca7ff8dd97b94 | compilade | git@compilade.net | 2025-07-02T13:10:24-04:00 | GitHub | noreply@github.com | 2025-07-02T13:10:24-04:00 | | llama : initial Mamba-2 support (#9126) |
| 1279 | e17991c466ac835b2c71bd813c1ca7ff8dd97b94 | c46944aa25b5560e6195f8cbcbc8947e9a6395c3 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T19:35:47+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T20:08:45+03:00 | | sync : ggml |
| 1280 | c46944aa25b5560e6195f8cbcbc8947e9a6395c3 | f3ed38d793fab9f97987dc52d7a41622f2056701 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-07-02T13:55:32+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T20:08:45+03:00 | | ggml : add version function to get lib version (ggml/1286) |
| 1281 | 55a1c5a5fdefef808e95aabd3d5563af1068cc80 | 12a81af45f0dbbab24bd819a15f57c03ceb1be90 | Aman Gupta | amangupta052@gmail.com | 2025-07-02T20:34:24+08:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T15:48:33+03:00 | | CUDA: add softmax broadcast (#14475) |
| 1282 | 12a81af45f0dbbab24bd819a15f57c03ceb1be90 | 8875523eb311cac832bfda0c581e852292185ae9 | Johannes Gäßler | johannesg@5d6.de | 2025-07-02T13:42:12+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T15:48:33+03:00 | | CUDA: broadcasting for FlashAttention mask (#14500) |
| 1283 | 8875523eb311cac832bfda0c581e852292185ae9 | ec68e84c32325a3417fbcd2e60d4bda6adb4e4bc | Jeff Bolz | jbolz@nvidia.com | 2025-07-01T03:32:56-05:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T15:48:33+03:00 | | vulkan: support softmax/FA batch and broadcast (#14449) |
| 1284 | ec68e84c32325a3417fbcd2e60d4bda6adb4e4bc | 307e79d33d4cdd9f1d6c42fc861724e6ba12b98f | Georgi Gerganov | ggerganov@gmail.com | 2025-06-27T21:50:57+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T15:48:33+03:00 | | ggml : support bcast ggml_soft_max_ext, ggml_flash_attn_ext (#14435) |
| 1285 | 307e79d33d4cdd9f1d6c42fc861724e6ba12b98f | d7f5f4e578d1f60b0835d1734a50438c309b3e5c | zhouwg | zhouwg2000@gmail.com | 2025-07-02T20:38:10+08:00 | GitHub | noreply@github.com | 2025-07-02T14:38:10+02:00 | | opencl : fix possible buffer overflow in dump_tensor (#14490) |
| 1286 | d7f5f4e578d1f60b0835d1734a50438c309b3e5c | c8a4e470f65c3a932e18eddc5fba1844876d7463 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-02T14:12:07+03:00 | GitHub | noreply@github.com | 2025-07-02T14:12:07+03:00 | | simple-chat : fix context-exceeded condition (#14494) |
| 1287 | c8a4e470f65c3a932e18eddc5fba1844876d7463 | 603e43dc913f90d9c3291624728b731687eddc35 | Eric Zhang | 34133756+EZForever@users.noreply.github.com | 2025-07-02T19:00:04+08:00 | GitHub | noreply@github.com | 2025-07-02T13:00:04+02:00 | | opencl : skip empty nodes on cgraph compute (#14491) |
| 1288 | 603e43dc913f90d9c3291624728b731687eddc35 | 611ba4b264c8fbb9d2d978bd6a5c17aeab1700fb | lhez | quic_lih@quicinc.com | 2025-07-02T00:07:42-07:00 | GitHub | noreply@github.com | 2025-07-02T09:07:42+02:00 | | opencl : update upscale to support align corners (#14488) |
| 1289 | 611ba4b264c8fbb9d2d978bd6a5c17aeab1700fb | 85841e121db5b96ae4d277a3cd3995eef66b98ec | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-02T09:02:51+02:00 | GitHub | noreply@github.com | 2025-07-02T09:02:51+02:00 | | ci : add OpenCL to labeler workflow (#14496) |
| 1290 | 85841e121db5b96ae4d277a3cd3995eef66b98ec | 68b3cd65141586ef13a698baa9029627f038f102 | Eric Zhang | 34133756+EZForever@users.noreply.github.com | 2025-07-02T13:41:35+08:00 | GitHub | noreply@github.com | 2025-07-02T08:41:35+03:00 | | github : add OpenCL backend to issue templates (#14492) |
| 1291 | 68b3cd65141586ef13a698baa9029627f038f102 | de569441470332ff922c23fb0413cc957be75b25 | Björn Ganster | mail@bjoern-ganster.de | 2025-07-02T07:19:31+02:00 | GitHub | noreply@github.com | 2025-07-02T08:19:31+03:00 | | ggml : Callback before abort (#14481) |
| 1292 | de569441470332ff922c23fb0413cc957be75b25 | 1b2aaf28acd632b79bc3b07acb5dad877dc7dbb1 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T18:04:08+03:00 | GitHub | noreply@github.com | 2025-07-01T18:04:08+03:00 | | ci : disable fast-math for Metal GHA CI (#14478) |
| 1293 | 1b2aaf28acd632b79bc3b07acb5dad877dc7dbb1 | 343b6e94b61674ac94e64ede7b7ea3365793f7ed | Grzegorz Grasza | xek@redhat.com | 2025-07-01T15:44:11+02:00 | GitHub | noreply@github.com | 2025-07-01T15:44:11+02:00 | | Add Vulkan images to docker.md (#14472) |
| 1294 | 343b6e94b61674ac94e64ede7b7ea3365793f7ed | 6a746cf9c40f8d93d13eb8a6c2b989a15e7fa61a | Chenguang Li | 757486878@qq.com | 2025-07-01T16:47:30+08:00 | GitHub | noreply@github.com | 2025-07-01T16:47:30+08:00 | | CANN: update aclnnGroupedMatmulV2 to aclnnGroupedMatmulV3 (#14411) |
| 1295 | 6a746cf9c40f8d93d13eb8a6c2b989a15e7fa61a | eff5e45443a248f89cc61e64786f38b68b4da489 | Jeff Bolz | jbolz@nvidia.com | 2025-07-01T03:43:08-05:00 | GitHub | noreply@github.com | 2025-07-01T10:43:08+02:00 | | vulkan: Split large mul_mat_id to fit in shared memory (#14451) |
| 1296 | eff5e45443a248f89cc61e64786f38b68b4da489 | a6a47958a1d9edc628cedc2f9320c84b59fe1f4f | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-07-01T10:14:21+02:00 | GitHub | noreply@github.com | 2025-07-01T10:14:21+02:00 | | add GELU_ERF (#14455) |
| 1297 | a6a47958a1d9edc628cedc2f9320c84b59fe1f4f | f61c05d4b1f8ad498a5cf7f0f57a40383aad563c | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T11:05:48+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T11:06:39+03:00 | | ggml : remove trailing whitespace (#0) |
| 1298 | f61c05d4b1f8ad498a5cf7f0f57a40383aad563c | 431b2c24f31bf5411786c378b409d0202df8d53c | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T10:27:52+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T11:06:39+03:00 | | sync : ggml |
| 1299 | 497be7c01dab243100f9c42e0a763a746e42d038 | 79b33b231774d5c39c8df018e9a276becae6d41a | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-06-24T06:10:16+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-07-01T11:06:39+03:00 | | ggml-quants : rename best_mad to best_error (ggml/1283) |
| 1300 | 79b33b231774d5c39c8df018e9a276becae6d41a | 0a5a3b5cdfd887cf0f8e09d9ff89dee130cfcdde | lhez | quic_lih@quicinc.com | 2025-07-01T00:19:16-07:00 | GitHub | noreply@github.com | 2025-07-01T09:19:16+02:00 | | opencl : add GEGLU, REGLU, SWIGLU (#14456) |
| 1301 | 0a5a3b5cdfd887cf0f8e09d9ff89dee130cfcdde | 745f11fed09ba2303f3f494efb12286d382608e5 | Aman Gupta | amangupta052@gmail.com | 2025-06-30T23:57:04+08:00 | GitHub | noreply@github.com | 2025-06-30T23:57:04+08:00 | | Add Conv2d for CPU (#14388) |
| 1302 | 745f11fed09ba2303f3f494efb12286d382608e5 | 5dd942de5922a22ec8446a4ad2203738dbcb9389 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-30T18:03:03+03:00 | GitHub | noreply@github.com | 2025-06-30T18:03:03+03:00 | | memory : correctly handle failure in apply() (#14438) |
| 1303 | 5dd942de5922a22ec8446a4ad2203738dbcb9389 | a7417f55945eb7c7a5ea6807e66564b8066a4e50 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-30T17:04:05+03:00 | GitHub | noreply@github.com | 2025-06-30T17:04:05+03:00 | | metal : disable fast-math for some cpy kernels (#14460) |
| 1304 | a7417f55945eb7c7a5ea6807e66564b8066a4e50 | eb3fa2913e0ace38912011fcb331624c679656f5 | Romain Biessy | romain.biessy@codeplay.com | 2025-06-30T14:52:02+02:00 | GitHub | noreply@github.com | 2025-06-30T14:52:02+02:00 | | ggml-cpu: sycl: Re-enable exp f16 (#14462) |
| 1305 | eb3fa2913e0ace38912011fcb331624c679656f5 | c839a2da1a0d4275c3a98f70c3a7627487573a46 | Diego Devesa | slarengh@gmail.com | 2025-06-30T03:43:15-07:00 | GitHub | noreply@github.com | 2025-06-30T12:43:15+02:00 | | test-backend-ops : disable llama test (#14461) |
| 1306 | c839a2da1a0d4275c3a98f70c3a7627487573a46 | e9b6350e61d592634263a14b3d77ecbf6c1fb096 | xiaobing318 | 71554036+xiaobing318@users.noreply.github.com | 2025-06-30T17:48:24+08:00 | GitHub | noreply@github.com | 2025-06-30T12:48:24+03:00 | | cmake : Remove redundant include path in CMakeLists.txt (#14452) |
| 1307 | e9b6350e61d592634263a14b3d77ecbf6c1fb096 | caf5681fcb47dfe9bafee94ef9aa8f669ac986c7 | Vedran Miletić | vedran@miletic.net | 2025-06-30T10:17:18+02:00 | GitHub | noreply@github.com | 2025-06-30T10:17:18+02:00 | | scripts : make the shell scripts cross-platform (#14341) |
| 1308 | caf5681fcb47dfe9bafee94ef9aa8f669ac986c7 | 83790b0e7e09ab17238b16452a33053a71dbdfad | matteo | matteo.serva@gmail.com | 2025-06-29T20:02:53+02:00 | GitHub | noreply@github.com | 2025-06-29T20:02:53+02:00 | | server : support jinja extra template kwargs (Qwen3 enable_thinking feature), from command line and from client (#13196) |
| 1309 | 83790b0e7e09ab17238b16452a33053a71dbdfad | f47c1d7106e49062279bcc57fc1077c0db61e278 | Renat | rntk@users.noreply.github.com | 2025-06-29T19:29:57+02:00 | GitHub | noreply@github.com | 2025-06-29T19:29:57+02:00 | | server : fix appearance of the chats list context menu for Safari (#14322) |
| 1310 | f47c1d7106e49062279bcc57fc1077c0db61e278 | a5d1fb6212298db1be1639db4c03adb2c522ee13 | Akarshan Biswas | akarshan@menlo.ai | 2025-06-29T21:07:58+05:30 | GitHub | noreply@github.com | 2025-06-29T21:07:58+05:30 | | SYCL: disable faulty fp16 exp kernel (#14395) |
| 1311 | a5d1fb6212298db1be1639db4c03adb2c522ee13 | a0535ffa0d35fccfec3e1a0a3bfc9dbb6054d7c0 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-29T14:38:10+02:00 | GitHub | noreply@github.com | 2025-06-29T14:38:10+02:00 | | ggml : fix unmerged GGML_FPxx_TO_FPxx refactoring (#14443) |
| 1312 | a0535ffa0d35fccfec3e1a0a3bfc9dbb6054d7c0 | bd9c981d7226107f18deb8344c3301450311bb8b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-29T11:04:10+02:00 | GitHub | noreply@github.com | 2025-06-29T11:04:10+02:00 | | ggml : implement REGLU/GEGLU/SWIGLU ops (#14158) |
| 1313 | bd9c981d7226107f18deb8344c3301450311bb8b | 27208bf657cfe7262791df473927225e48efe482 | Jeff Bolz | jbolz@nvidia.com | 2025-06-29T02:43:36-05:00 | GitHub | noreply@github.com | 2025-06-29T09:43:36+02:00 | | vulkan: Add fusion support for RMS_NORM+MUL (#14366) |
| 1314 | 27208bf657cfe7262791df473927225e48efe482 | 63a7bb3c7e1c6b0a92d03b0a594d3cd501d6ed3e | Aman Gupta | amangupta052@gmail.com | 2025-06-29T01:30:53+08:00 | GitHub | noreply@github.com | 2025-06-29T01:30:53+08:00 | | CUDA: add bf16 and f32 support to cublas_mul_mat_batched (#14361) |
| 1315 | 63a7bb3c7e1c6b0a92d03b0a594d3cd501d6ed3e | 00d5282c7f2a0bb05bb315fb81ca3b0f42cf9f07 | Jeff Bolz | jbolz@nvidia.com | 2025-06-28T10:36:40-05:00 | GitHub | noreply@github.com | 2025-06-28T17:36:40+02:00 | | vulkan: handle noncontig in the final case of ggml_vk_get_cpy_pipeline (#14378) |
| 1316 | 00d5282c7f2a0bb05bb315fb81ca3b0f42cf9f07 | 566c16fcce44876a167c37f159085afe6f84b28c | Jeff Bolz | jbolz@nvidia.com | 2025-06-28T10:17:09-05:00 | GitHub | noreply@github.com | 2025-06-28T17:17:09+02:00 | | vulkan: lock accesses of pinned_memory vector (#14333) |
| 1317 | 566c16fcce44876a167c37f159085afe6f84b28c | b25e92774e2fa4ee3820e458d5cf43f40190f8d2 | Weizhao Ouyang | weizhao.ouyang@arm.com | 2025-06-28T22:08:21+08:00 | GitHub | noreply@github.com | 2025-06-28T16:08:21+02:00 | | model : add support for ERNIE 4.5 0.3B model (#14408) |
| 1318 | b25e92774e2fa4ee3820e458d5cf43f40190f8d2 | 6609507a910aa7437aaa53fd999447de3947d998 | Xinpeng Dou | 15529241576@163.com | 2025-06-28T17:35:41+08:00 | GitHub | noreply@github.com | 2025-06-28T17:35:41+08:00 | | fix async_mode bug (#14432) |
| 1319 | 6609507a910aa7437aaa53fd999447de3947d998 | ceb1bf5a34d5e66e28b23dcc7a3cd83fe1e27481 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-28T09:57:07+02:00 | GitHub | noreply@github.com | 2025-06-28T09:57:07+02:00 | | ci : fix windows build and release (#14431) |
| 1320 | ceb1bf5a34d5e66e28b23dcc7a3cd83fe1e27481 | 72babea5dea56c8a8e8420ccf731b12a5cf37854 | Jeff Bolz | jbolz@nvidia.com | 2025-06-27T22:35:30-05:00 | GitHub | noreply@github.com | 2025-06-27T22:35:30-05:00 | | vulkan: Fix GGML_VULKAN_SHADER_DEBUG_INFO (#14427) |
| 1321 | 72babea5dea56c8a8e8420ccf731b12a5cf37854 | 43678060c1f4cfab2f899466fd615e358677f807 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-27T21:42:02+03:00 | GitHub | noreply@github.com | 2025-06-27T21:42:02+03:00 | | graph : make llm_graph_context destructor virtual (#14410) |
| 1322 | 43678060c1f4cfab2f899466fd615e358677f807 | 8d94219a4a7f2da72ee542019ca01f36af93d1d6 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-27T17:55:45+03:00 | GitHub | noreply@github.com | 2025-06-27T17:55:45+03:00 | | recurrent : call balloc split_reset() in init_batch() (#14414) |
| 1323 | 8d94219a4a7f2da72ee542019ca01f36af93d1d6 | f667f1e6244e1f420512fa66692b7096ff17f366 | Radoslav Gerganov | rgerganov@gmail.com | 2025-06-27T16:41:40+03:00 | GitHub | noreply@github.com | 2025-06-27T16:41:40+03:00 | | ggml : add ggml_set_rows (#14274) |
| 1324 | f667f1e6244e1f420512fa66692b7096ff17f366 | 8846aace4934ad29651ea61b8c7e3f6b0556e3d2 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-27T10:42:19+02:00 | GitHub | noreply@github.com | 2025-06-27T10:42:19+02:00 | | convert : fix broken sentencepiece vocab (#14416) |
| 1325 | 8846aace4934ad29651ea61b8c7e3f6b0556e3d2 | a01047b041aa04aeea351933658433ed004516ab | Xuan-Son Nguyen | son@huggingface.co | 2025-06-26T19:34:02+02:00 | GitHub | noreply@github.com | 2025-06-26T20:34:02+03:00 | | model : gemma3n text-only (#14400) |
| 1326 | a01047b041aa04aeea351933658433ed004516ab | b25346221dadb9101aa9dda55431dde4d3596943 | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-06-26T13:46:53-03:00 | GitHub | noreply@github.com | 2025-06-26T13:46:53-03:00 | | cmake: regen vulkan shaders when shaders-gen sources change (#14398) |
| 1327 | b25346221dadb9101aa9dda55431dde4d3596943 | e8215dbb96b8fb94a24c29cdd228166fb972dbfc | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-26T15:01:14+02:00 | GitHub | noreply@github.com | 2025-06-26T15:01:14+02:00 | | llama : return mistral-v7-tekken as default template only (#14390) |
| 1328 | e8215dbb96b8fb94a24c29cdd228166fb972dbfc | 5783ae43599400b723b5da0569c1f848419ff3c7 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-26T15:51:19+03:00 | GitHub | noreply@github.com | 2025-06-26T15:51:19+03:00 | | metal : add special-case mat-vec mul for ne00 == 4 (#14385) |
| 1329 | 5783ae43599400b723b5da0569c1f848419ff3c7 | bf5bcd0b857db420235e03639f0a5f218a7f8cf8 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-26T15:50:15+03:00 | GitHub | noreply@github.com | 2025-06-26T15:50:15+03:00 | | metal : batch rows copy in a single threadgroup (#14384) |
| 1330 | bf5bcd0b857db420235e03639f0a5f218a7f8cf8 | 716301d1b03c31875ec3b24526c48c8b1bd0fd8c | Aaron Teo | aaron.teo1@ibm.com | 2025-06-26T18:41:41+08:00 | GitHub | noreply@github.com | 2025-06-26T12:41:41+02:00 | | docs: update s390x documentation + add faq (#14389) |
| 1331 | ba67d6c03fed3193987a4ef13050e6e2a4451df4 | 8f52c6f8abb2ed9ad967cc0bfb90acb073011193 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-26T17:42:32+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-26T17:42:32+08:00 | | llama-calrt: | infer-op: copy data from output buffer to dst-tensor.data |
| 1332 | 8f52c6f8abb2ed9ad967cc0bfb90acb073011193 | d2f325130c4ed77a3dcff2c23a926587fb08e2cc | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-26T15:10:45+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-26T15:10:45+08:00 | | Squashed commit of the following: |
| 1333 | 716301d1b03c31875ec3b24526c48c8b1bd0fd8c | 60ef23d6c14d325d83eae5752e5de39ad268e9b0 | R0CKSTAR | yeahdongcn@gmail.com | 2025-06-26T12:11:59+08:00 | GitHub | noreply@github.com | 2025-06-26T12:11:59+08:00 | | musa: enable fp16 mma (all) and cublas on qy2 (#13842) |
| 1334 | 60ef23d6c14d325d83eae5752e5de39ad268e9b0 | b193d5306912a2adae0fde7481819f6ee0941bc6 | Aaron Teo | aaron.teo1@ibm.com | 2025-06-26T05:49:04+08:00 | GitHub | noreply@github.com | 2025-06-25T23:49:04+02:00 | | ggml-cpu: enable IBM NNPA Vector Intrinsics (#14317) |
| 1335 | b193d5306912a2adae0fde7481819f6ee0941bc6 | 2bf9d539dd158345e3a3b096e16474af535265b4 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-25T23:26:51+02:00 | GitHub | noreply@github.com | 2025-06-25T23:26:51+02:00 | | ggml : do not output unprintable characters on GGUF load failure (#14381) |
| 1336 | 2bf9d539dd158345e3a3b096e16474af535265b4 | 73e53dc834c0a2336cd104473af6897197b96277 | Anton Mitkov | anton_b_mitkov@abv.bg | 2025-06-25T17:09:55+01:00 | GitHub | noreply@github.com | 2025-06-25T18:09:55+02:00 | | sycl: GGML_SYCL_DISABLE_OPT on by default for all Intel Devices (#13973) |
| 1337 | 73e53dc834c0a2336cd104473af6897197b96277 | 62af464227dafa1c55e0535bcb24346326748f46 | lhez | quic_lih@quicinc.com | 2025-06-24T11:46:25-07:00 | GitHub | noreply@github.com | 2025-06-24T11:46:25-07:00 | | opencl: ref count `ggml_backend_opencl_context` and refactor profiling (#14254) |
| 1338 | 62af464227dafa1c55e0535bcb24346326748f46 | c148cf1946275952a79ad50b6199725f12a70411 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-24T18:26:30+03:00 | GitHub | noreply@github.com | 2025-06-24T18:26:30+03:00 | | batch : fix check for empty sequences in memory (#14364) |
| 1339 | c148cf1946275952a79ad50b6199725f12a70411 | 1b809cee225222094a0ff5be8467240487ce4ae4 | Mathieu Baudier | mbaudier@argeo.org | 2025-06-24T15:05:31+02:00 | GitHub | noreply@github.com | 2025-06-24T15:05:31+02:00 | | cmake : use LLAMA_BUILD_NUMBER when defining LLAMA_INSTALL_VERSION (#14362) |
| 1340 | 1b809cee225222094a0ff5be8467240487ce4ae4 | abf241045d09cad70dc797b0fba393ad09ee2cbe | Nigel Bosch | pnigelb@gmail.com | 2025-06-24T08:59:11Z | GitHub | noreply@github.com | 2025-06-24T10:59:11+02:00 | | server : move no API key doc to /health (#14352) |
| 1341 | abf241045d09cad70dc797b0fba393ad09ee2cbe | 901e20bbe571fbde48d13eb188f4e7cdc7562fb6 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-24T09:31:00+02:00 | GitHub | noreply@github.com | 2025-06-24T09:31:00+02:00 | | main : honor --verbose-prompt on interactive prompts (#14350) |
| 1342 | 901e20bbe571fbde48d13eb188f4e7cdc7562fb6 | 0142961a2e67909e33cdf410274b56c08c5dce7a | Bartowski | 3266127+bartowski1182@users.noreply.github.com | 2025-06-24T02:17:58-04:00 | GitHub | noreply@github.com | 2025-06-24T09:17:58+03:00 | | jinja : Add Mistral-Small-3.2-24B-Instruct-2506.jinja (#14349) |
| 1343 | 0142961a2e67909e33cdf410274b56c08c5dce7a | ce82bd0117bd3598300b3a089d13d401b90279c7 | uvos | philipp@uvos.xyz | 2025-06-24T01:12:56+02:00 | GitHub | noreply@github.com | 2025-06-24T01:12:56+02:00 | | CUDA/HIP: optimize mmv paths taken for HIP devices (#14324) |
| 1344 | ce82bd0117bd3598300b3a089d13d401b90279c7 | bf2a99e3cb06a14dce2a586450d012dc31f922ae | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-06-23T15:30:51-03:00 | GitHub | noreply@github.com | 2025-06-23T15:30:51-03:00 | | ci: add workflow for relocatable cmake package (#14346) |
| 1345 | bf2a99e3cb06a14dce2a586450d012dc31f922ae | 72c6bc3f3d0cf3bf160dddf1b803fed52bcbb0a3 | Jeff Bolz | jbolz@nvidia.com | 2025-06-23T08:44:48-05:00 | GitHub | noreply@github.com | 2025-06-23T15:44:48+02:00 | | vulkan: update windows SDK in release.yml (#14344) |
| 1346 | 72c6bc3f3d0cf3bf160dddf1b803fed52bcbb0a3 | defe2158dd5e250b4ef53994057a4ec03131a263 | Molly Sophia | mollysophia379@gmail.com | 2025-06-23T19:56:19+08:00 | GitHub | noreply@github.com | 2025-06-23T19:56:19+08:00 | | llama : better rwkv chat template and add missing `inputs.use_jinja` setting (#14336) |
| 1347 | defe2158dd5e250b4ef53994057a4ec03131a263 | 7b50d589a863c7631135c1226f6eab65cb406212 | Johannes Gäßler | johannesg@5d6.de | 2025-06-23T13:11:31+02:00 | GitHub | noreply@github.com | 2025-06-23T13:11:31+02:00 | | CUDA: mul_mat_v support for batch sizes > 1 (#14262) |
| 1348 | 7b50d589a863c7631135c1226f6eab65cb406212 | 3a9457df962b5883f2773f1c295e8c19df60d89f | Georgi Gerganov | ggerganov@gmail.com | 2025-06-23T12:27:35+03:00 | GitHub | noreply@github.com | 2025-06-23T12:27:35+03:00 | | kv-cells : fix tracking of seq_pos (#14339) |
| 1349 | 3a9457df962b5883f2773f1c295e8c19df60d89f | fa4a9f2a1ccda2573189a9d4995bdf0bceb41156 | Jeff Bolz | jbolz@nvidia.com | 2025-06-23T03:19:24-05:00 | GitHub | noreply@github.com | 2025-06-23T10:19:24+02:00 | | vulkan: update windows SDK in CI (#14334) |
| 1350 | fa4a9f2a1ccda2573189a9d4995bdf0bceb41156 | 238005c2dc67426cf678baa2d54c881701693288 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-06-22T22:16:26+01:00 | GitHub | noreply@github.com | 2025-06-22T23:16:26+02:00 | | quantize : handle user-defined pruning of whole layers (blocks) (#13037) |
| 1351 | 238005c2dc67426cf678baa2d54c881701693288 | 66aba7aca9a245d71f1ddf02c7a97223529752a8 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-22T19:46:17+02:00 | GitHub | noreply@github.com | 2025-06-22T19:46:17+02:00 | | gguf-py : fix SpecialVocab parsing when post_processor is null (#14330) |
| 1352 | 66aba7aca9a245d71f1ddf02c7a97223529752a8 | f1f5e82df6222dcbaca6396c0de44df259a1694f | Ruikai Peng | retr0@retr0.blog | 2025-06-23T01:28:06+08:00 | GitHub | noreply@github.com | 2025-06-23T01:28:06+08:00 | | run : avoid double tokenization (#14327) |
| 1353 | f1f5e82df6222dcbaca6396c0de44df259a1694f | af3373f1adfca56119f3e4de0e6a0a8df8edf3d9 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-22T20:10:07+03:00 | GitHub | noreply@github.com | 2025-06-22T20:10:07+03:00 | | examples : fix is_first logic for tokenization (#14329) |
| 1354 | af3373f1adfca56119f3e4de0e6a0a8df8edf3d9 | 5d5c066de8a3d2cb32f04c4d5ad1560945f30bf3 | uvos | philipp@uvos.xyz | 2025-06-22T16:51:23+02:00 | GitHub | noreply@github.com | 2025-06-22T16:51:23+02:00 | | HIP: enable vec fattn on RDNA4 (#14323) |
| 1355 | 5d5c066de8a3d2cb32f04c4d5ad1560945f30bf3 | 40bfa04c95c19fb42bafd4e21b5c2a7771846801 | yuiseki | yuiseki@gmail.com | 2025-06-22T21:44:57+09:00 | GitHub | noreply@github.com | 2025-06-22T14:44:57+02:00 | | mtmd : fix Pixtral OOM with large images by capping image_size to 1024 (#14326) |
| 1356 | 40bfa04c95c19fb42bafd4e21b5c2a7771846801 | aa064b2eb7f7779db5e094a9d8d66d5033557de2 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-22T07:37:43+02:00 | GitHub | noreply@github.com | 2025-06-22T08:37:43+03:00 | | common : use std::string_view now that we target c++17 (#14319) |
| 1357 | aa064b2eb7f7779db5e094a9d8d66d5033557de2 | aa0ef5c578eef4c2adc7be1282f21bab5f3e8d26 | Aman Gupta | amangupta052@gmail.com | 2025-06-22T12:39:54+08:00 | GitHub | noreply@github.com | 2025-06-22T12:39:54+08:00 | | CUDA: add mean operation (#14313) |
| 1358 | aa0ef5c578eef4c2adc7be1282f21bab5f3e8d26 | bb16041caef45cd4348cd6f84906b5dfec7a1f6a | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-21T18:12:05+02:00 | GitHub | noreply@github.com | 2025-06-21T18:12:05+02:00 | | gguf-py : fix Qwen3-Embedding eos token (#14314) |
| 1359 | bb16041caef45cd4348cd6f84906b5dfec7a1f6a | 58cba76a9aab728717509d62b67c13afd8dc227a | Markus Tavenrath | mtavenrath@users.noreply.github.com | 2025-06-21T08:17:12+02:00 | GitHub | noreply@github.com | 2025-06-21T08:17:12+02:00 | | Add support for VK_EXT_debug_utils to add labels to Vulkan objects. (#13792) |
| 1360 | 58cba76a9aab728717509d62b67c13afd8dc227a | 67ae5312e255ae97852a4a216e2245580bfafd72 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-21T07:33:21+02:00 | GitHub | noreply@github.com | 2025-06-21T07:33:21+02:00 | | gguf-py : fix TemplateProcessing pair when bos/eos is missing (#14312) |
| 1361 | 67ae5312e255ae97852a4a216e2245580bfafd72 | 692e3cdd0a069ab56411b64506a67537d767683e | Georgi Gerganov | ggerganov@gmail.com | 2025-06-21T08:04:18+03:00 | GitHub | noreply@github.com | 2025-06-21T08:04:18+03:00 | | metal : fix thread-safety (#14300) |
| 1362 | 692e3cdd0a069ab56411b64506a67537d767683e | b23fa0b3f40165ca3aae8ad4ee756e72f9a130dd | Georgi Gerganov | ggerganov@gmail.com | 2025-06-21T08:03:46+03:00 | GitHub | noreply@github.com | 2025-06-21T08:03:46+03:00 | | memory : rename interface to llama_memory_context_i (#14296) |
| 1363 | b23fa0b3f40165ca3aae8ad4ee756e72f9a130dd | 06cbedfca1587473df9b537f1dd4d6bfa2e3de13 | Daniel Han | danielhanchen@gmail.com | 2025-06-20T21:32:01-07:00 | GitHub | noreply@github.com | 2025-06-21T06:32:01+02:00 | | convert : fix Llama 4 conversion (#14311) |
| 1364 | 06cbedfca1587473df9b537f1dd4d6bfa2e3de13 | b7147673f26c7ceb926a43c4734002bba291bcb8 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T20:50:24+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T21:02:47+03:00 | | sync : ggml |
| 1365 | b7147673f26c7ceb926a43c4734002bba291bcb8 | d860dd99a4178f58d1d1fa64eebc2aabc95392a7 | Acly | aclysia@gmail.com | 2025-06-18T13:34:50+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T21:02:47+03:00 | | Add `ggml_roll` (ggml/1274) |
| 1366 | d860dd99a4178f58d1d1fa64eebc2aabc95392a7 | c959f462a0b4d42eaf930ffd72df0e435c97d5d5 | David Chiu | david20571015@gmail.com | 2025-06-21T01:43:35+08:00 | GitHub | noreply@github.com | 2025-06-20T19:43:35+02:00 | | docs : fix the link to llama.h (#14293) |
| 1367 | c959f462a0b4d42eaf930ffd72df0e435c97d5d5 | 22015b2092e291022ea3cfedc6aeb8f2643807da | Aman Gupta | amangupta052@gmail.com | 2025-06-20T22:48:24+08:00 | GitHub | noreply@github.com | 2025-06-20T22:48:24+08:00 | | CUDA: add conv_2d_transpose (#14287) |
| 1368 | 22015b2092e291022ea3cfedc6aeb8f2643807da | dd6e6d0b6a4bbe3ebfc931d1eb14db2f2b1d70af | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-20T16:37:44+02:00 | GitHub | noreply@github.com | 2025-06-20T16:37:44+02:00 | | lint : remove trailing whitepace (#14304) |
| 1369 | dd6e6d0b6a4bbe3ebfc931d1eb14db2f2b1d70af | 8308f98c7fb778e54bf75538f5234d8bd20915e9 | Ruikai Peng | retr0@retr0.blog | 2025-06-20T22:13:06+08:00 | GitHub | noreply@github.com | 2025-06-20T07:13:06-07:00 | | vocab : prevent tokenizer overflow (#14301) |
| 1370 | 8308f98c7fb778e54bf75538f5234d8bd20915e9 | 6369be07359d03723f38c0a4a014ff0f698a0738 | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-06-20T15:07:21+02:00 | GitHub | noreply@github.com | 2025-06-20T15:07:21+02:00 | | sycl: add usage of enqueue_functions extension (#14244) |
| 1371 | 6369be07359d03723f38c0a4a014ff0f698a0738 | 88fc854b4bd2e3caf10e705e6afcbbca136f0a3c | Christian Kastner | ckk@kvr.at | 2025-06-20T12:17:32Z | GitHub | noreply@github.com | 2025-06-20T14:17:32+02:00 | | Implement GGML_CPU_ALL_VARIANTS for PowerPC (#14286) |
| 1372 | 88fc854b4bd2e3caf10e705e6afcbbca136f0a3c | e28c1b93fd7d3f8faf9551d962e8a65fe2122e38 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-20T14:04:09+02:00 | GitHub | noreply@github.com | 2025-06-20T14:04:09+02:00 | | llama : improve sep token handling (#14272) |
| 1373 | e28c1b93fd7d3f8faf9551d962e8a65fe2122e38 | d27b3ca1758dfb1718e333d497ef4b68ad109bc2 | Diego Devesa | slarengh@gmail.com | 2025-06-20T04:57:36-07:00 | GitHub | noreply@github.com | 2025-06-20T13:57:36+02:00 | | cuda : synchronize graph capture and cublas handle destruction (#14288) |
| 1374 | d27b3ca1758dfb1718e333d497ef4b68ad109bc2 | 9230dbe2c757e2d5071329095727d0fa9d4b85c4 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T11:19:15+03:00 | GitHub | noreply@github.com | 2025-06-20T11:19:15+03:00 | | ggml : fix repack work size for mul_mat_id (#14292) |
| 1375 | 9230dbe2c757e2d5071329095727d0fa9d4b85c4 | 812939a9e90f99d1bd5bb1bc6b99d12600671d50 | Charles Xu | charles.xu@arm.com | 2025-06-20T09:51:01+02:00 | GitHub | noreply@github.com | 2025-06-20T10:51:01+03:00 | | ggml: Update KleidiAI to v1.9.0 (#14277) |
| 1376 | 812939a9e90f99d1bd5bb1bc6b99d12600671d50 | 4c9fdfbe1580a66fd7d77c77418ce2c606a29fdd | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T10:50:27+03:00 | GitHub | noreply@github.com | 2025-06-20T10:50:27+03:00 | | model : more uniform output id handling (#14275) |
| 1377 | 4c9fdfbe1580a66fd7d77c77418ce2c606a29fdd | 9eaa51e7f08593f123f00136591179a8f5956ecd | Georgi Gerganov | ggerganov@gmail.com | 2025-06-20T10:14:14+03:00 | GitHub | noreply@github.com | 2025-06-20T10:14:14+03:00 | | ubatch : new splitting logic (#14217) |
| 1378 | d2f325130c4ed77a3dcff2c23a926587fb08e2cc | 6391ee7ffd97e723dfbdcc32f001e1b95a63cba6 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-17T11:53:33+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-20T09:57:18+08:00 | | calrt: add input_token in graph; input_token will be copy to calrt_in_buffer |
| 1379 | 9eaa51e7f08593f123f00136591179a8f5956ecd | 8f71d0f3e86ccbba059350058af8758cafed73e6 | Aman Gupta | amangupta052@gmail.com | 2025-06-20T09:50:24+08:00 | GitHub | noreply@github.com | 2025-06-20T09:50:24+08:00 | | CUDA: add conv_2d_dw (#14265) |
| 1380 | 8f71d0f3e86ccbba059350058af8758cafed73e6 | 381174bbdaf10d6a80dc2099f284b20544d86962 | Diego Devesa | slarengh@gmail.com | 2025-06-19T12:24:14-07:00 | GitHub | noreply@github.com | 2025-06-19T21:24:14+02:00 | | ggml-cpu : remove unnecesary arm feature detection (#14281) |
| 1381 | 381174bbdaf10d6a80dc2099f284b20544d86962 | d67341dc18fc5cc63362880ab2f8f9ecfc7932e7 | Alex Trotta | 44127594+Ahajha@users.noreply.github.com | 2025-06-19T09:56:12-04:00 | GitHub | noreply@github.com | 2025-06-19T15:56:12+02:00 | | gguf-py : make sentencepiece optional (#14200) |
| 1382 | d67341dc18fc5cc63362880ab2f8f9ecfc7932e7 | 456af35eb70177b8dd5779b6d4c21bb020f9cebd | aa956 | aa956@users.noreply.github.com | 2025-06-19T16:01:03+03:00 | GitHub | noreply@github.com | 2025-06-19T16:01:03+03:00 | | server : add server parameters for draft model cache type (#13782) |
| 1383 | 456af35eb70177b8dd5779b6d4c21bb020f9cebd | 600e3e9b50c1f0c9fc4a70356241fd87f00e8e14 | fanyang | fanyang89@outlook.com | 2025-06-19T20:49:48+08:00 | GitHub | noreply@github.com | 2025-06-19T14:49:48+02:00 | | build : suppress gcc15 compile warnings (#14261) |
| 1384 | 600e3e9b50c1f0c9fc4a70356241fd87f00e8e14 | fffcce535ebbdc25f81966a15b758658788e7466 | Anton Mitkov | anton_b_mitkov@abv.bg | 2025-06-19T11:40:21+01:00 | GitHub | noreply@github.com | 2025-06-19T11:40:21+01:00 | | sycl: Cleanup codepaths in Get Rows in sycl backend (#14215) |
| 1385 | fffcce535ebbdc25f81966a15b758658788e7466 | 5fc7856815920f828f9e90cb759ac82c7f0c1ea5 | bashayer hijji | bashayer.hijji@gmail.com | 2025-06-19T13:24:12+03:00 | GitHub | noreply@github.com | 2025-06-19T12:24:12+02:00 | | llama-bench : add --no-warmup flag (#14224) (#14270) |
| 1386 | 5fc7856815920f828f9e90cb759ac82c7f0c1ea5 | faed5a5f5dde4d816adecdccb4f4682a04126f92 | pqnet | 119850+pqnet@users.noreply.github.com | 2025-06-19T12:21:40+02:00 | GitHub | noreply@github.com | 2025-06-19T12:21:40+02:00 | | convert : fix remote option in Windows (#14100) |
| 1387 | faed5a5f5dde4d816adecdccb4f4682a04126f92 | 10bb545c5b54175ed9874ad8d187effa2bcb4b5f | Aaron Teo | aaron.teo1@ibm.com | 2025-06-19T17:48:54+08:00 | GitHub | noreply@github.com | 2025-06-19T11:48:54+02:00 | | llamafile : support s390x SIMD instruction set (#14273) |
| 1388 | 10bb545c5b54175ed9874ad8d187effa2bcb4b5f | edc4a29effe716956fdfd2bc5b9cba2a6d8492f8 | 0cc4m | picard12@live.de | 2025-06-19T09:15:42+02:00 | GitHub | noreply@github.com | 2025-06-19T09:15:42+02:00 | | Vulkan: Set device max size for host memory to avoid OOM warning and fallback to CPU buffer (#14249) |
| 1389 | edc4a29effe716956fdfd2bc5b9cba2a6d8492f8 | ed3290ab34493fd3d2e0f925f816a401da4f2dfd | Gabe Goodhart | gabe.l.hart@gmail.com | 2025-06-19T00:08:14-05:00 | GitHub | noreply@github.com | 2025-06-19T08:08:14+03:00 | | memory : Hybrid recurrent cache (#13979) |
| 1390 | ed3290ab34493fd3d2e0f925f816a401da4f2dfd | 8d947136546773f6410756f37fcc5d3e65b8135d | Georgi Gerganov | ggerganov@gmail.com | 2025-06-19T08:05:21+03:00 | GitHub | noreply@github.com | 2025-06-19T08:05:21+03:00 | | metal : add mean kernel (#14267) |
| 1391 | 8d947136546773f6410756f37fcc5d3e65b8135d | 50d2227953ca9024f04255b4f116d06fcc0db74c | Aaron Teo | aaron.teo1@ibm.com | 2025-06-19T01:10:26+08:00 | GitHub | noreply@github.com | 2025-06-18T18:10:26+01:00 | | docs: add s390x build documentation (#14264) |
| 1392 | 50d2227953ca9024f04255b4f116d06fcc0db74c | 6231c5cd6d49d61511f328b5f43322407df90a91 | Aaron Teo | aaron.teo1@ibm.com | 2025-06-19T01:10:08+08:00 | GitHub | noreply@github.com | 2025-06-18T18:10:08+01:00 | | ggml-cpu: reduce asm calls for hsum (#14037) |
| 1393 | 6231c5cd6d49d61511f328b5f43322407df90a91 | ef035803eb9dbc306ea9a8ff82e30af12b567cf7 | Aaron Teo | aaron.teo1@ibm.com | 2025-06-19T01:06:49+08:00 | GitHub | noreply@github.com | 2025-06-18T18:06:49+01:00 | | ggml-cpu: fix uncaught underscore terminators (#14023) |
| 1394 | ef035803eb9dbc306ea9a8ff82e30af12b567cf7 | 413977de32e90712ecec84d0b9c738847da8dc02 | Charles Xu | charles.xu@arm.com | 2025-06-18T13:40:07+02:00 | GitHub | noreply@github.com | 2025-06-18T12:40:07+01:00 | | ggml: Add Apple support for GGML_CPU_ALL_VARIANTS (#14258) |
| 1395 | 413977de32e90712ecec84d0b9c738847da8dc02 | 95402553a5effc61ddc9e29c7bcb56f71311dd4a | Xuan-Son Nguyen | son@huggingface.co | 2025-06-18T10:43:57+02:00 | GitHub | noreply@github.com | 2025-06-18T10:43:57+02:00 | | mtmd : refactor llava-uhd preprocessing logic (#14247) |
| 1396 | 95402553a5effc61ddc9e29c7bcb56f71311dd4a | 3865cff4f5b84c16119590efd6ce537789a27715 | Xuan-Son Nguyen | son@huggingface.co | 2025-06-18T09:58:43+02:00 | GitHub | noreply@github.com | 2025-06-18T09:58:43+02:00 | | llama-chat : fix multiple system message for gemma, orion (#14246) |
| 1397 | 3865cff4f5b84c16119590efd6ce537789a27715 | d03172cc797b6adbbc00bfc09caf614fb0f895a0 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-18T09:52:07+02:00 | GitHub | noreply@github.com | 2025-06-18T09:52:07+02:00 | | convert : fix null head_dim AutoConfig regression (#14248) |
| 1398 | d03172cc797b6adbbc00bfc09caf614fb0f895a0 | dd8e59f4435342eda93c5b0cf4109e21c9c7d0eb | Georgi Gerganov | ggerganov@gmail.com | 2025-06-18T09:58:23+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-18T09:59:21+03:00 | | sync : ggml |
| 1399 | dd8e59f4435342eda93c5b0cf4109e21c9c7d0eb | bbe98d27840453c8787d18470963530fdc27d89f | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-06-13T15:06:42+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-18T09:59:21+03:00 | | ggml : disable warnings for tests when using MSVC (ggml/1273) |
| 1400 | bbe98d27840453c8787d18470963530fdc27d89f | c2056ed6d461e6d5432f04f221e221ab795dc652 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-06-13T09:05:44+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-18T09:59:21+03:00 | | ggml : remove unused ggml_context_container (ggml/1272) |
| 1401 | c2056ed6d461e6d5432f04f221e221ab795dc652 | c46503014db0d63fa7b1b28c58adfb51054e2dec | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-06-12T12:27:09+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-18T09:59:21+03:00 | | examples : include examples in msvc disable warn (ggml/1270) |
| 1402 | c46503014db0d63fa7b1b28c58adfb51054e2dec | 860a9e4eeff3eb2e7bd1cc38f65787cc6c8177af | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-06-17T17:33:25-03:00 | GitHub | noreply@github.com | 2025-06-17T22:33:25+02:00 | | cmake: remove shader-gen step-targets from ggml-vulkan (#14226) |
| 1403 | 860a9e4eeff3eb2e7bd1cc38f65787cc6c8177af | fe9d60e74a6cb71bcaed2029377bfa2872b4abb0 | xctan | xc-tan@outlook.com | 2025-06-17T17:58:32+08:00 | GitHub | noreply@github.com | 2025-06-17T12:58:32+03:00 | | ggml-cpu : remove the weak alias trick (#14221) |
| 1404 | fe9d60e74a6cb71bcaed2029377bfa2872b4abb0 | e434e69183fd9e1031f4445002083178c331a28b | R0CKSTAR | yeahdongcn@gmail.com | 2025-06-17T17:48:08+08:00 | GitHub | noreply@github.com | 2025-06-17T17:48:08+08:00 | | musa: fix build warning (unused variable) (#14231) |
| 1405 | e434e69183fd9e1031f4445002083178c331a28b | 89fea80d298184d1cd93564f48e060d9f541f4b4 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-16T21:58:42+02:00 | GitHub | noreply@github.com | 2025-06-16T21:58:42+02:00 | | common : suggest --jinja when autodetection fails (#14222) |
| 1406 | 89fea80d298184d1cd93564f48e060d9f541f4b4 | 6adc3c3ebc029af058ac950a8e2a825fdf18ecc6 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-16T22:33:27+03:00 | GitHub | noreply@github.com | 2025-06-16T22:33:27+03:00 | | server : fix incorrect usage of llama_get_embeddings() (#14225) |
| 1407 | 6adc3c3ebc029af058ac950a8e2a825fdf18ecc6 | 0dbcabde8c006d5cf781ca0fe071c41559572a72 | Diego Devesa | slarengh@gmail.com | 2025-06-16T08:11:43-07:00 | GitHub | noreply@github.com | 2025-06-16T08:11:43-07:00 | | llama : add thread safety test (#14035) |
| 1408 | 0dbcabde8c006d5cf781ca0fe071c41559572a72 | ad590be98c83217fcf1a101d78d9ab389fd5dc0b | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-06-16T10:32:13-03:00 | GitHub | noreply@github.com | 2025-06-16T10:32:13-03:00 | | cmake: clean up external project logic for vulkan-shaders-gen (#14179) |
| 1409 | ad590be98c83217fcf1a101d78d9ab389fd5dc0b | 7d6d91babfa129906b39c9099eca4234c44f4f1e | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-06-16T21:53:41+09:00 | GitHub | noreply@github.com | 2025-06-16T14:53:41+02:00 | | model : add NeoBERT (#14164) |
| 1410 | 7d6d91babfa129906b39c9099eca4234c44f4f1e | d3e64b9f490cee41b7b9aa275dae2f6568ae3054 | uvos | philipp@uvos.xyz | 2025-06-16T13:47:38+02:00 | GitHub | noreply@github.com | 2025-06-16T13:47:38+02:00 | | HIP: disable rocwmma on gfx12 by default until rocm 7.0 (#14202) |
| 1411 | d3e64b9f490cee41b7b9aa275dae2f6568ae3054 | 3ba0d843c6bd3faea5cf5e53dc7f3c82be20bffb | Georgi Gerganov | ggerganov@gmail.com | 2025-06-16T14:14:00+03:00 | GitHub | noreply@github.com | 2025-06-16T14:14:00+03:00 | | llama : rework embeddings logic (#14208) |
| 1412 | 3ba0d843c6bd3faea5cf5e53dc7f3c82be20bffb | 0bf49eb668bb95b50e41583e22aaf60ddade1fbe | Charles Xu | charles.xu@arm.com | 2025-06-16T11:47:57+02:00 | GitHub | noreply@github.com | 2025-06-16T11:47:57+02:00 | | ggml: Add Android support for GGML_CPU_ALL_VARIANTS (#14206) |
| 1413 | 0bf49eb668bb95b50e41583e22aaf60ddade1fbe | 4ad243677bca6c97f14dbc187b2116b51fcb7ffd | Bartowski | 3266127+bartowski1182@users.noreply.github.com | 2025-06-16T09:16:06+01:00 | GitHub | noreply@github.com | 2025-06-16T10:16:06+02:00 | | convert : remove arcee change in convert_hf_to_gguf_update.py (#14207) |
| 1414 | 4ad243677bca6c97f14dbc187b2116b51fcb7ffd | c89c2d1ab94b11845240b7d3313c87691ea18d88 | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-06-16T16:20:59+09:00 | GitHub | noreply@github.com | 2025-06-16T09:20:59+02:00 | | gguf-py : allow key override when adding value to GGUFWriter (#14194) |
| 1415 | c89c2d1ab94b11845240b7d3313c87691ea18d88 | 3555b3004ba7687be3d734acade52a3345758aa4 | Jeff Bolz | jbolz@nvidia.com | 2025-06-16T00:21:08-06:00 | GitHub | noreply@github.com | 2025-06-16T08:21:08+02:00 | | vulkan: mutex around vkQueueSubmit (#14127) |
| 1416 | 3555b3004ba7687be3d734acade52a3345758aa4 | d7da8dc83a03b30e1ec10317080082ea76840c38 | xctan | xc-tan@outlook.com | 2025-06-16T13:54:15+08:00 | GitHub | noreply@github.com | 2025-06-16T13:54:15+08:00 | | ggml-cpu : rework weak alias on apple targets (#14146) |
| 1417 | d7da8dc83a03b30e1ec10317080082ea76840c38 | cd355eda7df1898d25d433b4bdaa4b4b479e0bad | Bartowski | 3266127+bartowski1182@users.noreply.github.com | 2025-06-16T00:04:06+01:00 | GitHub | noreply@github.com | 2025-06-16T01:04:06+02:00 | | model : Add support for Arcee AI's upcoming AFM model (#14185) |
| 1418 | cd355eda7df1898d25d433b4bdaa4b4b479e0bad | 30e5b01de2a0bcddc7c063c8ef0802703a958417 | Eric Curtin | ecurtin@redhat.com | 2025-06-15T23:36:22+02:00 | GitHub | noreply@github.com | 2025-06-15T23:36:22+02:00 | | server : When listening on a unix domain socket don't print http:// and port (#14180) |
| 1419 | 30e5b01de2a0bcddc7c063c8ef0802703a958417 | e54b394082de242be4ee2e692b11fcc8d4eba371 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-06-15T17:53:45+01:00 | GitHub | noreply@github.com | 2025-06-15T18:53:45+02:00 | | quantize : change int to unsigned int for KV overrides (#14197) |
| 1420 | e54b394082de242be4ee2e692b11fcc8d4eba371 | 2c2caa444341d99c87ff153f142c2d4762a776a2 | uvos | philipp@uvos.xyz | 2025-06-15T17:30:13+02:00 | GitHub | noreply@github.com | 2025-06-15T17:30:13+02:00 | | CUDA/HIP: fix ssm_scan on devices where warp size is not 32 (#14196) |
| 1421 | 2c2caa444341d99c87ff153f142c2d4762a776a2 | 5fce5f948df8f189a5401a8ecaa9753106e75abb | uvos | philipp@uvos.xyz | 2025-06-15T15:45:27+02:00 | GitHub | noreply@github.com | 2025-06-15T15:45:27+02:00 | | HIP: Replace usage of depricated preprocessor macro __AMDGCN_WAVEFRONT_SIZE__ (#14183) |
| 1422 | 5fce5f948df8f189a5401a8ecaa9753106e75abb | 9ae4143bc6ecb4c2f0f0301578f619f6c201b857 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-15T10:52:11+03:00 | GitHub | noreply@github.com | 2025-06-15T10:52:11+03:00 | | kv-cache : fix use-after-move of defrag info (#14189) |
| 1423 | 9ae4143bc6ecb4c2f0f0301578f619f6c201b857 | c311ac664d68d10781a3e7b9f02d9d9520837d80 | Mikko Juola | mikjuo@gmail.com | 2025-06-15T00:52:06-07:00 | GitHub | noreply@github.com | 2025-06-15T09:52:06+02:00 | | model : add dots.llm1 architecture support (#14044) (#14118) |
| 1424 | c311ac664d68d10781a3e7b9f02d9d9520837d80 | b9912ac570de8945ae9383c9ca8291027bf287dd | Georgi Gerganov | ggerganov@gmail.com | 2025-06-15T10:08:58+03:00 | GitHub | noreply@github.com | 2025-06-15T10:08:58+03:00 | | cparams : rename LLAMA_MAX_PARALLEL_SEQUENCES to LLAMA_MAX_SEQ (#14188) |
| 1425 | b9912ac570de8945ae9383c9ca8291027bf287dd | 00ba7726100d7e1941d9f5a06f56a7559945b33c | Georgi Gerganov | ggerganov@gmail.com | 2025-06-15T09:18:37+03:00 | GitHub | noreply@github.com | 2025-06-15T09:18:37+03:00 | | batch : auto-gen positions + verify multi-sequence input (#14177) |
| 1426 | 00ba7726100d7e1941d9f5a06f56a7559945b33c | 3cb203c89f60483e349f841684173446ed23c28f | Pepijn de Vos | me@pepijndevos.nl | 2025-06-15T08:06:37+02:00 | GitHub | noreply@github.com | 2025-06-15T08:06:37+02:00 | | docs : remove WIP since PR has been merged (#13912) |
| 1427 | 3cb203c89f60483e349f841684173446ed23c28f | 2e42be42bd6bf1dcc643d6ac4e77419bfe5dd24f | Piotr | piotr.stankiewicz@docker.com | 2025-06-14T18:25:15+02:00 | GitHub | noreply@github.com | 2025-06-14T17:25:15+01:00 | | llama-chat : Do not throw when tool parsing fails (#14012) |
| 1428 | 2e42be42bd6bf1dcc643d6ac4e77419bfe5dd24f | fb85a288d72abbd5e5daa8de96e6f8bfa7b5ab46 | Aman Gupta | amangupta052@gmail.com | 2025-06-14T16:34:20+08:00 | GitHub | noreply@github.com | 2025-06-14T10:34:20+02:00 | | compare-llama-bench: add option to plot (#14169) |
| 1429 | fb85a288d72abbd5e5daa8de96e6f8bfa7b5ab46 | 40643edb86eb10b471b0f57d4f3f7eb0e06a0df7 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T20:03:05+03:00 | GitHub | noreply@github.com | 2025-06-13T20:03:05+03:00 | | vocab : fix build (#14175) |
| 1430 | 40643edb86eb10b471b0f57d4f3f7eb0e06a0df7 | 3cfbbdb44e08fd19429fed6cc85b982a91f0efd5 | Svetlozar Georgiev | 55534064+sgeor255@users.noreply.github.com | 2025-06-13T17:32:56+01:00 | GitHub | noreply@github.com | 2025-06-13T18:32:56+02:00 | | sycl: fix docker image (#14144) |
| 1431 | 3cfbbdb44e08fd19429fed6cc85b982a91f0efd5 | 80709b70a2f87c13ccaf1480b799393109996789 | Guy Goldenberg | guy110698@gmail.com | 2025-06-13T19:20:25+03:00 | GitHub | noreply@github.com | 2025-06-13T19:20:25+03:00 | | Merge commit from fork |
| 1432 | 80709b70a2f87c13ccaf1480b799393109996789 | 26ff3685bfbaa4c8838d7afd988b17dd5eb99f92 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T18:35:00+03:00 | GitHub | noreply@github.com | 2025-06-13T18:35:00+03:00 | | batch : add LLAMA_BATCH_DEBUG environment variable (#14172) |
| 1433 | 26ff3685bfbaa4c8838d7afd988b17dd5eb99f92 | 60c666347becacff81cd4bc9a52038ba71038e41 | ddpasa | 112642920+ddpasa@users.noreply.github.com | 2025-06-13T15:17:53+02:00 | GitHub | noreply@github.com | 2025-06-13T15:17:53+02:00 | | docs : Update multimodal.md (#14122) |
| 1434 | 60c666347becacff81cd4bc9a52038ba71038e41 | b7cc7745e38141060845145af0a1fd489d8e3e33 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T13:47:55+03:00 | GitHub | noreply@github.com | 2025-06-13T13:47:55+03:00 | | batch : rework llama_batch_allocr (#14153) |
| 1435 | b7cc7745e38141060845145af0a1fd489d8e3e33 | cc8d08187918c6f643c3ffabb7b1ac21aa19f3d1 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T11:55:44+03:00 | GitHub | noreply@github.com | 2025-06-13T11:55:44+03:00 | | readme : remove survey link (#14168) |
| 1436 | cc8d08187918c6f643c3ffabb7b1ac21aa19f3d1 | d714dadb57d8feaa03d13b79345a4c3382172d61 | Christian Kastner | ckk@kvr.at | 2025-06-13T08:38:52Z | GitHub | noreply@github.com | 2025-06-13T10:38:52+02:00 | | cmake: Add ability to pass in LLAMA_BUILD_NUMBER/COMMIT (#14167) |
| 1437 | d714dadb57d8feaa03d13b79345a4c3382172d61 | ffad04397399ea1650fda6560c7c753059804876 | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-06-13T17:34:08+09:00 | GitHub | noreply@github.com | 2025-06-13T11:34:08+03:00 | | pooling : make cls_b and cls_out_b optional (#14165) |
| 1438 | ffad04397399ea1650fda6560c7c753059804876 | 0889eba570126f8a2f5a0e88fde776bbc91cca66 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T11:18:25+03:00 | GitHub | noreply@github.com | 2025-06-13T11:18:25+03:00 | | server : fix SWA condition for full context reprocess (#14163) |
| 1439 | 0889eba570126f8a2f5a0e88fde776bbc91cca66 | c61285e7396c8e526fe7794c19e8d4f1c99bfc51 | Anton Mitkov | anton.mitkov@codeplay.com | 2025-06-13T08:51:39+01:00 | GitHub | noreply@github.com | 2025-06-13T08:51:39+01:00 | | sycl: Adding additional cpy dbg print output (#14034) |
| 1440 | c61285e7396c8e526fe7794c19e8d4f1c99bfc51 | 09cf2c7c655c90e53e100f29b830a788bab0653d | Ewan Crawford | ewan@codeplay.com | 2025-06-13T08:45:37+01:00 | GitHub | noreply@github.com | 2025-06-13T08:45:37+01:00 | | SYCL: Bump oneMath commit (#14152) |
| 1441 | 09cf2c7c655c90e53e100f29b830a788bab0653d | c33fe8b8c4427202706b1434e9fc8ab5752c9cac | Christian Kastner | ckk@kvr.at | 2025-06-13T06:51:34Z | GitHub | noreply@github.com | 2025-06-13T09:51:34+03:00 | | cmake : Improve build-info.cpp generation (#14156) |
| 1442 | c33fe8b8c4427202706b1434e9fc8ab5752c9cac | ed52f3668e633423054a4eab61bb7efee47025ab | Georgi Gerganov | ggerganov@gmail.com | 2025-06-13T08:03:54+03:00 | GitHub | noreply@github.com | 2025-06-13T08:03:54+03:00 | | vocab : prevent heap overflow when vocab is too small (#14145) |
| 1443 | 6391ee7ffd97e723dfbdcc32f001e1b95a63cba6 | b9850050a5b5147cb2e6720b98ed8c1e8e02eff9 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-13T10:58:06+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-13T10:58:06+08:00 | | update .gitignore to update cal_test.gguf |
| 1444 | b9850050a5b5147cb2e6720b98ed8c1e8e02eff9 | 933b205c32551848589b8339437d778cb4f2a46d | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-06T17:25:52+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-13T09:20:23+08:00 | | calrt: add call chain for using calrt::infer |
| 1445 | ed52f3668e633423054a4eab61bb7efee47025ab | a681b4ba83a61dce71a4f24e558efe7278d8b1a9 | Anton Mitkov | anton.mitkov@codeplay.com | 2025-06-12T14:15:11+01:00 | GitHub | noreply@github.com | 2025-06-12T15:15:11+02:00 | | sycl: Remove not needed copy f16->f32 for dnnl mul mat (#14125) |
| 1446 | a681b4ba83a61dce71a4f24e558efe7278d8b1a9 | 7d516443dd7766569110b38e6374649bee6eb1c4 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T14:43:09+03:00 | GitHub | noreply@github.com | 2025-06-12T14:43:09+03:00 | | readme : remove project status link (#14149) |
| 1447 | 933b205c32551848589b8339437d778cb4f2a46d | 33fea9a8592fd338abe279b89350a3af2e0ff1f7 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-12T17:47:19+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-12T17:50:26+08:00 | | gguf-py: add cal_test.py |
| 1448 | 7d516443dd7766569110b38e6374649bee6eb1c4 | f6e1a7aa8787b5c00acba6370cb70a0beff48b1e | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T11:51:38+03:00 | GitHub | noreply@github.com | 2025-06-12T11:51:38+03:00 | | server : re-enable SWA speculative decoding (#14131) |
| 1449 | f6e1a7aa8787b5c00acba6370cb70a0beff48b1e | c3ee46fab49a765d2e32e171e9ed7a5fa121dd9c | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T11:50:01+03:00 | GitHub | noreply@github.com | 2025-06-12T11:50:01+03:00 | | context : simplify output counting logic during decode (#14142) |
| 1450 | c3ee46fab49a765d2e32e171e9ed7a5fa121dd9c | e2c0b6e46a5596665569ae765f0993cea2619af6 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T11:49:26+03:00 | GitHub | noreply@github.com | 2025-06-12T11:49:26+03:00 | | batch : remove logits_all flag (#14141) |
| 1451 | e2c0b6e46a5596665569ae765f0993cea2619af6 | 9596506965f65be5d802ecef6a315fe43d2391a8 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T10:14:24+03:00 | GitHub | noreply@github.com | 2025-06-12T10:14:24+03:00 | | cmake : handle whitepsaces in path during metal build (#14126) |
| 1452 | 9596506965f65be5d802ecef6a315fe43d2391a8 | a20b2b05bce6622c585459ebf46f142f113d021c | Georgi Gerganov | ggerganov@gmail.com | 2025-06-12T10:02:15+03:00 | GitHub | noreply@github.com | 2025-06-12T10:02:15+03:00 | | kv-cache : fix split_equal handling in unified implementation (#14130) |
| 1453 | a20b2b05bce6622c585459ebf46f142f113d021c | 2e89f76b7af2c0b827be785e445f2e2b3e52e1ca | compilade | git@compilade.net | 2025-06-12T02:56:04-04:00 | GitHub | noreply@github.com | 2025-06-12T02:56:04-04:00 | | context : round n_tokens to next multiple of n_seqs when reserving (#14140) |
| 1454 | 2e89f76b7af2c0b827be785e445f2e2b3e52e1ca | 532802f938c6a18cc6a704057ab571f253fd77ed | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-06-11T17:19:44-03:00 | GitHub | noreply@github.com | 2025-06-11T17:19:44-03:00 | | common: fix issue with regex_escape routine on windows (#14133) |
| 1455 | 532802f938c6a18cc6a704057ab571f253fd77ed | d4e0d95cf581f50c9a21d06eaecae2dd580076bd | Christian Kastner | ckk@kvr.at | 2025-06-11T19:07:44Z | GitHub | noreply@github.com | 2025-06-11T21:07:44+02:00 | | Implement GGML_CPU_ALL_VARIANTS for ARM (#14080) |
| 1456 | d4e0d95cf581f50c9a21d06eaecae2dd580076bd | cc66a7f78f95dcb5208420f4dd1abc3fc6aec0cc | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-11T19:04:23+02:00 | GitHub | noreply@github.com | 2025-06-11T19:04:23+02:00 | | chore : clean up relative source dir paths (#14128) |
| 1457 | cc66a7f78f95dcb5208420f4dd1abc3fc6aec0cc | bd248d4dc7265437e1918dc53ae3f49a0c592e5f | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-11T17:16:32+02:00 | GitHub | noreply@github.com | 2025-06-11T17:16:32+02:00 | | tests : add test-tokenizers-repo (#14017) |
| 1458 | bd248d4dc7265437e1918dc53ae3f49a0c592e5f | 7781e5fe99f2a4fc1c8af0a8488eedac4644cb72 | Jeff Bolz | jbolz@nvidia.com | 2025-06-11T09:48:52-05:00 | GitHub | noreply@github.com | 2025-06-11T09:48:52-05:00 | | vulkan: Better thread-safety for command pools/buffers (#14116) |
| 1459 | 7781e5fe99f2a4fc1c8af0a8488eedac4644cb72 | 89a184fa7125b7c3e2fb567337cc2bd7bc2d376c | Aman | amangupta052@gmail.com | 2025-06-11T22:42:25+08:00 | GitHub | noreply@github.com | 2025-06-11T16:42:25+02:00 | | webui: Wrap long numbers instead of infinite horizontal scroll (#14062) |
| 1460 | 89a184fa7125b7c3e2fb567337cc2bd7bc2d376c | 2baf07727f921d9a4a1b63a2eff941e95d0488ed | Georgi Gerganov | ggerganov@gmail.com | 2025-06-11T16:48:45+03:00 | GitHub | noreply@github.com | 2025-06-11T16:48:45+03:00 | | kv-cache : relax SWA masking condition (#14119) |
| 1461 | 2baf07727f921d9a4a1b63a2eff941e95d0488ed | 7ae2932116a4de3141cbc8c488f7f6e4e06d8171 | Taylor | quantumtraveling@gmail.com | 2025-06-11T06:43:43-04:00 | GitHub | noreply@github.com | 2025-06-11T13:43:43+03:00 | | server : pass default --keep argument (#14120) |
| 1462 | 7ae2932116a4de3141cbc8c488f7f6e4e06d8171 | 1f7d50b2936023b26eb218e944e62834b80a2ce0 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-11T12:52:45+03:00 | GitHub | noreply@github.com | 2025-06-11T12:52:45+03:00 | | kv-cache : add LLAMA_KV_CACHE_DEBUG environment variable (#14121) |
| 1463 | 1f7d50b2936023b26eb218e944e62834b80a2ce0 | 4c763c8d1b4d4de20bf364ec1837430783cba984 | Jeff Bolz | jbolz@nvidia.com | 2025-06-11T00:19:25-05:00 | GitHub | noreply@github.com | 2025-06-11T07:19:25+02:00 | | vulkan: Track descriptor pools/sets per-context (#14109) |
| 1464 | 4c763c8d1b4d4de20bf364ec1837430783cba984 | dad5c44398b78467943ed0a603ea427fe9f6fc62 | lhez | quic_lih@quicinc.com | 2025-06-10T16:55:58-07:00 | GitHub | noreply@github.com | 2025-06-10T16:55:58-07:00 | | opencl: add `mul_mv_id_q4_0_f32_8x_flat` (#14003) |
| 1465 | dad5c44398b78467943ed0a603ea427fe9f6fc62 | 55f6b9fa6563f6ae49113f9abdc980c12348cc1c | compilade | git@compilade.net | 2025-06-10T18:20:14-04:00 | GitHub | noreply@github.com | 2025-06-10T18:20:14-04:00 | | kv-cache : avoid modifying recurrent cells when setting inputs (#13834) |
| 1466 | 55f6b9fa6563f6ae49113f9abdc980c12348cc1c | 3678b838bb71eaccbaeb479ff38c2e12bfd2f960 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-10T23:29:52+02:00 | GitHub | noreply@github.com | 2025-06-10T23:29:52+02:00 | | convert : fix duplicate key DeepSeek-R1 conversion error (#14103) |
| 1467 | 3678b838bb71eaccbaeb479ff38c2e12bfd2f960 | 652b70e6678d503626f3c9d33831ba604b473401 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-10T18:02:08+02:00 | GitHub | noreply@github.com | 2025-06-10T18:02:08+02:00 | | llama : support GEGLU for jina-bert-v2 (#14090) |
| 1468 | 652b70e6678d503626f3c9d33831ba604b473401 | 3a12db23b6918682ccb70ab15d897f11fe5fc320 | Jeff Bolz | jbolz@nvidia.com | 2025-06-10T10:53:47-05:00 | GitHub | noreply@github.com | 2025-06-10T10:53:47-05:00 | | vulkan: force device 0 in CI (#14106) |
| 1469 | 3a12db23b6918682ccb70ab15d897f11fe5fc320 | ae92c1855b1a5b604fa1bfcced19a556ab3e78c5 | Juk Armstrong | 69222624+jukofyork@users.noreply.github.com | 2025-06-10T16:48:07+01:00 | GitHub | noreply@github.com | 2025-06-10T16:48:07+01:00 | | Fixed spec timings to: accepted/tested instead of accepted/drafted (#14104) |
| 1470 | ae92c1855b1a5b604fa1bfcced19a556ab3e78c5 | b7ce1ad1e332da5909772d173a7e6748fbd1887a | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T17:37:45+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T18:39:33+03:00 | | sync : ggml |
| 1471 | b7ce1ad1e332da5909772d173a7e6748fbd1887a | 97340b4c9924be86704dbf155e97c8319849ee19 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T11:34:10+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T18:39:33+03:00 | | ggml : fix weak alias win32 (whisper/0) |
| 1472 | 97340b4c9924be86704dbf155e97c8319849ee19 | 2bb0467043258bdc58dbaefb33786f1731b38937 | 0cc4m | picard12@live.de | 2025-06-10T14:01:33+02:00 | GitHub | noreply@github.com | 2025-06-10T13:01:33+01:00 | | Vulkan: Don't default to CPU device (like llvmpipe), even if no other device is available, to allow fallback to CPU backend (#14099) |
| 1473 | 2bb0467043258bdc58dbaefb33786f1731b38937 | b8e2194efc529378be45ab9b27d6648a5b81458a | Isaac McFadyen | isaac@imcf.me | 2025-06-10T02:41:01-04:00 | GitHub | noreply@github.com | 2025-06-10T09:41:01+03:00 | | rpc : nicer error messages for RPC server crash (#14076) |
| 1474 | b8e2194efc529378be45ab9b27d6648a5b81458a | 1a3b5e80f77c51e6d9fb7762f19b9e61fceca3de | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T09:20:51+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T09:21:56+03:00 | | sync : ggml |
| 1475 | 1a3b5e80f77c51e6d9fb7762f19b9e61fceca3de | 1f63e75f3b5dc7f44dbe63c8a41d23958fe95bc0 | Kai Pastor | dg0yt@darc.de | 2025-06-03T12:33:28+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-10T09:21:56+03:00 | | Add in-build ggml::ggml ALIAS library (ggml/1260) |
| 1476 | 1f63e75f3b5dc7f44dbe63c8a41d23958fe95bc0 | 40cbf571c9210031340f022d104317e89a1a6bb2 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-09T23:05:02+03:00 | GitHub | noreply@github.com | 2025-06-09T23:05:02+03:00 | | metal : use less stack memory in FA kernel (#14088) |
| 1477 | 40cbf571c9210031340f022d104317e89a1a6bb2 | 7f4fbe5183b23b6b2e25fd1ccc5d1fa8bb010cb7 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-09T23:04:35+03:00 | GitHub | noreply@github.com | 2025-06-09T23:04:35+03:00 | | kv-cache : fix shift and defrag logic (#14081) |
| 1478 | 7f4fbe5183b23b6b2e25fd1ccc5d1fa8bb010cb7 | f470bc36bed4d836b9de5a483fa0dfaee176d6f5 | Diego Devesa | slarengh@gmail.com | 2025-06-09T11:03:09-07:00 | GitHub | noreply@github.com | 2025-06-09T20:03:09+02:00 | | llama : allow building all tests on windows when not using shared libs (#13980) |
| 1479 | f470bc36bed4d836b9de5a483fa0dfaee176d6f5 | 8f47e25f56e9792093b7497c68e9f80bab82ed19 | xctan | axunlei@gmail.com | 2025-06-09T22:47:13+08:00 | GitHub | noreply@github.com | 2025-06-09T16:47:13+02:00 | | ggml-cpu : split arch-specific implementations (#13892) |
| 1480 | 8f47e25f56e9792093b7497c68e9f80bab82ed19 | 201b31dc2e6fef3de598280b53fd3b69420a4e49 | Diego Devesa | slarengh@gmail.com | 2025-06-09T07:36:26-07:00 | GitHub | noreply@github.com | 2025-06-09T16:36:26+02:00 | | cuda : fix device sync on buffer clear (#14033) |
| 1481 | 201b31dc2e6fef3de598280b53fd3b69420a4e49 | e21d2d4ae26d6c5daf1602a1061aa4f8e722ca57 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-09T17:17:31+03:00 | GitHub | noreply@github.com | 2025-06-09T17:17:31+03:00 | | graph : fix geglu (#14077) |
| 1482 | e21d2d4ae26d6c5daf1602a1061aa4f8e722ca57 | dc0623fddb661926e0998538895155e4f081ff09 | Xinpeng Dou | 15529241576@163.com | 2025-06-09T19:47:39+08:00 | GitHub | noreply@github.com | 2025-06-09T19:47:39+08:00 | | CANN: Simplify the environment variable setting(#13104) |
| 1483 | dc0623fddb661926e0998538895155e4f081ff09 | 87d34b381d5868e75586210ec17b5ef5deddc276 | R0CKSTAR | yeahdongcn@gmail.com | 2025-06-09T18:01:17+08:00 | GitHub | noreply@github.com | 2025-06-09T12:01:17+02:00 | | webui: fix sidebar being covered by main content (#14082) |
| 1484 | 87d34b381d5868e75586210ec17b5ef5deddc276 | b460d16ae858c6624fd37aec316622a4060ca325 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-09T12:57:58+03:00 | GitHub | noreply@github.com | 2025-06-09T12:57:58+03:00 | | server : fix LRU check (#14079) |
| 1485 | b460d16ae858c6624fd37aec316622a4060ca325 | 91a8ee6a6f1f4c8547ff7b745ef95c6edc1d2af6 | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-06-09T11:47:07+02:00 | GitHub | noreply@github.com | 2025-06-09T11:47:07+02:00 | | sycl: Add reorder to Q6_K mmvq implementation (#13885) |
| 1486 | 91a8ee6a6f1f4c8547ff7b745ef95c6edc1d2af6 | 056eb74534ffd3efc50a9b854156eb6876a52a44 | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-06-09T13:15:31+09:00 | GitHub | noreply@github.com | 2025-06-09T05:15:31+01:00 | | add geglu activation function (#14074) |
| 1487 | 056eb74534ffd3efc50a9b854156eb6876a52a44 | 247e5c6e447707bb4539bdf1913d206088a8fc69 | Yuanhao Ji | jiyuanhao@apache.org | 2025-06-09T11:20:06+08:00 | GitHub | noreply@github.com | 2025-06-09T11:20:06+08:00 | | CANN: Enable labeler for Ascend NPU (#13914) |
| 1488 | 247e5c6e447707bb4539bdf1913d206088a8fc69 | 5787b5da57e54dba760c2deeac1edf892e8fc450 | Diego Devesa | slarengh@gmail.com | 2025-06-08T11:39:56-07:00 | GitHub | noreply@github.com | 2025-06-08T11:39:56-07:00 | | cuda : fix buffer type check with integrated GPUs (#14069) |
| 1489 | 5787b5da57e54dba760c2deeac1edf892e8fc450 | 228f34c9ceefa3ea4f4d6933edd858121e8106cb | 吴小白 | 296015668@qq.com | 2025-06-07T21:39:11+08:00 | GitHub | noreply@github.com | 2025-06-07T10:39:11-03:00 | | ci: add LoongArch cross-compile build (#13944) |
| 1490 | 228f34c9ceefa3ea4f4d6933edd858121e8106cb | 0974ad7a7cd4bca846b15c484ff3be890135a52c | Akarshan Biswas | akarshan@menlo.ai | 2025-06-07T18:58:20+05:30 | GitHub | noreply@github.com | 2025-06-07T18:58:20+05:30 | | SYCL: Implement few same quantized type copy kernels (#13739) |
| 1491 | 0974ad7a7cd4bca846b15c484ff3be890135a52c | 745aa5319b9930068aff5e87cf5e9eef7227339b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-07T14:13:12+02:00 | GitHub | noreply@github.com | 2025-06-07T14:13:12+02:00 | | llama : fix llama_model_chat_template with template name (LLM_KV with suffix) (#14050) |
| 1492 | 745aa5319b9930068aff5e87cf5e9eef7227339b | 487a5e0401423bba02cd6e97e4d45131bb20b22b | Georgi Gerganov | ggerganov@gmail.com | 2025-06-06T14:11:15+03:00 | GitHub | noreply@github.com | 2025-06-06T14:11:15+03:00 | | llama : deprecate llama_kv_self_ API (#14030) |
| 1493 | 487a5e0401423bba02cd6e97e4d45131bb20b22b | d17a809ef0af09b16625e991a76f6fe80d9c332e | Georgi Gerganov | ggerganov@gmail.com | 2025-06-06T13:29:18+03:00 | GitHub | noreply@github.com | 2025-06-06T13:29:18+03:00 | | context : fix SWA-related warning for multiple sequences (#14045) |
| 1494 | 33fea9a8592fd338abe279b89350a3af2e0ff1f7 | 806f05a353d322e97d97e02f910de2b408a8b4ab | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-06T17:25:52+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-06T17:35:48+08:00 | | calrt: add load calbin |
| 1495 | d17a809ef0af09b16625e991a76f6fe80d9c332e | 1caae7fc6c77551cb1066515e0f414713eebb367 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-06T09:03:25+02:00 | GitHub | noreply@github.com | 2025-06-06T09:03:25+02:00 | | llama : support multiple classifier outputs and labels (#13940) |
| 1496 | 806f05a353d322e97d97e02f910de2b408a8b4ab | ec9e0301fef6476df83e94842c3b625501c95566 | Yunzhe Jia | yunzhe@calculet.tech | 2025-05-30T09:45:37+08:00 | Yunzhe Jia | yunzhe@calculet.tech | 2025-06-06T09:24:26+08:00 | | add calrt demo |
| 1497 | 1caae7fc6c77551cb1066515e0f414713eebb367 | 669c13e0f67edfba891c5bd7105eb4376bcf8626 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-05T17:42:31+02:00 | GitHub | noreply@github.com | 2025-06-05T17:42:31+02:00 | | gguf-py : add add_classifier_output_labels method to writer (#14031) |
| 1498 | 669c13e0f67edfba891c5bd7105eb4376bcf8626 | 146b88e8b3e16d22cea22f8be982a4d45e913b1e | Masato Nakasaka | rillomas@gmail.com | 2025-06-05T23:00:29+09:00 | GitHub | noreply@github.com | 2025-06-05T16:00:29+02:00 | | vulkan: Enable VK_KHR_cooperative_matrix extension for Intel Xe2 GPUs (#14001) |
| 1499 | 146b88e8b3e16d22cea22f8be982a4d45e913b1e | 7f37b6cf1e2c1b90bf0d9c8d91904b4b6c512748 | pockers21 | 134406831+pockers21@users.noreply.github.com | 2025-06-05T06:25:29-07:00 | GitHub | noreply@github.com | 2025-06-05T16:25:29+03:00 | | ci: fix CUDA build failure on autodl cloud machines (#14005) |
| 1500 | 7f37b6cf1e2c1b90bf0d9c8d91904b4b6c512748 | 3a077146a4761fdbd24bdd8eb098f46b8adc4dda | Georgi Gerganov | ggerganov@gmail.com | 2025-06-05T15:29:22+03:00 | GitHub | noreply@github.com | 2025-06-05T15:29:22+03:00 | | memory : migrate from llama_kv_cache to more generic llama_memory (#14006) |
| 1501 | 3a077146a4761fdbd24bdd8eb098f46b8adc4dda | d01d112abb055549d33bb8aac1755eb0c40918ad | Diego Devesa | slarengh@gmail.com | 2025-06-05T02:57:42-07:00 | GitHub | noreply@github.com | 2025-06-05T11:57:42+02:00 | | llama : allow using mmap without PrefetchVirtualMemory, apply GGML_WIN_VER to llama.cpp sources (#14013) |
| 1502 | d01d112abb055549d33bb8aac1755eb0c40918ad | 9f47fa5792bae5312615439e68dbf1826913f7ac | Olexandr88 | radole1203@gmail.com | 2025-06-05T10:50:55+03:00 | GitHub | noreply@github.com | 2025-06-05T10:50:55+03:00 | | readme : add badge (#13938) |
| 1503 | 9f47fa5792bae5312615439e68dbf1826913f7ac | 9e31bec4fd53634c9e5b04650488a09a055f5dab | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-05T09:29:18+02:00 | GitHub | noreply@github.com | 2025-06-05T09:29:18+02:00 | | vocab : warn about missing mask token (#14022) |
| 1504 | 9e31bec4fd53634c9e5b04650488a09a055f5dab | 5a8ae3053ced350ed300ba91600519fcad1c6ba7 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-05T09:06:29+03:00 | GitHub | noreply@github.com | 2025-06-05T09:06:29+03:00 | | context : fix pos_min initialization upon error decode (#14008) |
| 1505 | 5a8ae3053ced350ed300ba91600519fcad1c6ba7 | 0d3984424f2973c49c4bcabe4cc0153b4f90c601 | Jeff Bolz | jbolz@nvidia.com | 2025-06-05T00:17:58-05:00 | GitHub | noreply@github.com | 2025-06-05T07:17:58+02:00 | | vulkan: automatically deduce size of push constants (#13936) |
| 1506 | 0d3984424f2973c49c4bcabe4cc0153b4f90c601 | 3e63a58ef7addec35408e2eb67850d7cdc935dc3 | Ervin Áron Tasnádi | etasnadi@protonmail.com | 2025-06-04T22:02:00+02:00 | GitHub | noreply@github.com | 2025-06-04T22:02:00+02:00 | | ggml-vulkan: adds support for op CONV_TRANSPOSE_1D (#13813) |
| 1507 | 3e63a58ef7addec35408e2eb67850d7cdc935dc3 | 2589ad3704559f4dd860f5f303b19349c688a28a | Georgi Gerganov | ggerganov@gmail.com | 2025-06-04T18:58:20+03:00 | GitHub | noreply@github.com | 2025-06-04T18:58:20+03:00 | | kv-cache : refactor the update/defrag mechanism (#13988) |
| 1508 | 2589ad3704559f4dd860f5f303b19349c688a28a | 482548716f664f76e325ded58c9e8b7563e5e23a | Diego Devesa | slarengh@gmail.com | 2025-06-04T06:37:40-07:00 | GitHub | noreply@github.com | 2025-06-04T15:37:40+02:00 | | ci : remove cuda 11.7 releases, switch runner to windows 2022 (#13997) |
| 1509 | 482548716f664f76e325ded58c9e8b7563e5e23a | 3ac67535c86e2fc43e4eddf594412acc370bbb04 | Diego Devesa | slarengh@gmail.com | 2025-06-04T04:15:54-07:00 | GitHub | noreply@github.com | 2025-06-04T13:15:54+02:00 | | releases : use dl backend for linux release, remove arm64 linux release (#13996) |
| 1510 | 3ac67535c86e2fc43e4eddf594412acc370bbb04 | 0b4be4c435849b00dbd98b109cf7a22298d27b69 | Xuan-Son Nguyen | son@huggingface.co | 2025-06-04T10:11:26+02:00 | GitHub | noreply@github.com | 2025-06-04T10:11:26+02:00 | | llama-graph : use ggml_repeat_4d (#13998) |
| 1511 | 0b4be4c435849b00dbd98b109cf7a22298d27b69 | e0e806f52ebcd0ee285c994fe8fd8b8787d2cb0a | Johannes Gäßler | johannesg@5d6.de | 2025-06-04T08:57:05+02:00 | GitHub | noreply@github.com | 2025-06-04T08:57:05+02:00 | | CUDA: fix FTZ in FA for Gemma 3 (#13991) |
| 1512 | e0e806f52ebcd0ee285c994fe8fd8b8787d2cb0a | 7e00e60ef86645a01fda738fef85b74afa016a34 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-04T09:50:32+03:00 | GitHub | noreply@github.com | 2025-06-04T09:50:32+03:00 | | kv-cache : fix unified::seq_rm to work with seq_id < 0 (#13985) |
| 1513 | 7e00e60ef86645a01fda738fef85b74afa016a34 | ea1431b0fa3a8108aac1e0a94a13ccc4a749963e | Jeff Bolz | jbolz@nvidia.com | 2025-06-03T13:30:22-05:00 | GitHub | noreply@github.com | 2025-06-03T20:30:22+02:00 | | vulkan: fix warnings in perf logger querypool code (#13937) |
| 1514 | 71e74a3ac929b8af91f16f73f3c2b9b2f796d207 | bfb1e012a0b7658e8f00ed4333d059943ea9d648 | lhez | quic_lih@quicinc.com | 2025-06-02T16:54:58-07:00 | GitHub | noreply@github.com | 2025-06-02T16:54:58-07:00 | | opencl: add `backend_synchronize` (#13939) |
| 1515 | bfb1e012a0b7658e8f00ed4333d059943ea9d648 | 363757628848a27a435bbf22ff9476e9aeda5f40 | rmatif | 66360289+rmatif@users.noreply.github.com | 2025-06-02T23:53:36Z | GitHub | noreply@github.com | 2025-06-02T16:53:36-07:00 | | OpenCL: Add concat, tsembd, upscale, tanh, pad and repeat (#13840) |
| 1516 | 363757628848a27a435bbf22ff9476e9aeda5f40 | ea394d7ab1f8101716d48ce9421c94c71b93a00f | Georgi Gerganov | ggerganov@gmail.com | 2025-06-02T21:34:40+03:00 | GitHub | noreply@github.com | 2025-06-02T21:34:40+03:00 | | server : disable speculative decoding for SWA models (#13970) |
| 1517 | ea394d7ab1f8101716d48ce9421c94c71b93a00f | 5582c49c3961269eca96822abfb87528e942dd07 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-02T21:33:40+03:00 | GitHub | noreply@github.com | 2025-06-02T21:33:40+03:00 | | metal : use F32 accumulators in FA kernels (#13975) |
| 1518 | 5582c49c3961269eca96822abfb87528e942dd07 | c9bbc77931d223ed7e7cbcf1cb057bc02fd0db19 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-02T20:54:26+03:00 | GitHub | noreply@github.com | 2025-06-02T20:54:26+03:00 | | gemma : more consistent attention scaling for v2 and v3 (#13951) |
| 1519 | c9bbc77931d223ed7e7cbcf1cb057bc02fd0db19 | bfd322796cd838f906535ff3352624fc46338894 | Olivier Chafik | olivier.chafik@gmail.com | 2025-06-02T10:15:44-07:00 | GitHub | noreply@github.com | 2025-06-02T10:15:44-07:00 | | `server`: update deepseek reasoning format (pass reasoning_content as diffs) (#13933) |
| 1520 | bfd322796cd838f906535ff3352624fc46338894 | 093e3f1feb16e25e58f7d61e01266c830dd424b8 | Xuan-Son Nguyen | son@huggingface.co | 2025-06-02T16:29:28+02:00 | GitHub | noreply@github.com | 2025-06-02T16:29:28+02:00 | | mtmd : fix memory leak in mtmd_helper_eval_chunk_single (#13961) |
| 1521 | 093e3f1feb16e25e58f7d61e01266c830dd424b8 | 663445b0deb21fb602176da030d4154197a4fca6 | shalinib-ibm | Shalini.Salomi.Bodapati@ibm.com | 2025-06-02T17:48:36+05:30 | GitHub | noreply@github.com | 2025-06-02T15:18:36+03:00 | | cmake : Handle mixed-case 'Power' strings in POWER CPU detection (#13966) |
| 1522 | 663445b0deb21fb602176da030d4154197a4fca6 | 7675c555a13c9f473249e59a54db35032ce8e0fc | Atharva Dubey | atharva.dubey@codeplay.com | 2025-06-02T10:12:20+01:00 | GitHub | noreply@github.com | 2025-06-02T10:12:20+01:00 | | sycl: quantize and reorder the input to q8_1 when reorder is enabled (#13826) |
| 1523 | 7675c555a13c9f473249e59a54db35032ce8e0fc | 5e1c3aed4074480f63e914d6c44c93536ed1452a | Johannes Gäßler | johannesg@5d6.de | 2025-06-01T18:08:05+02:00 | GitHub | noreply@github.com | 2025-06-01T18:08:05+02:00 | | gguf: fix failure on version == 0 (#13956) |
| 1524 | 5e1c3aed4074480f63e914d6c44c93536ed1452a | c496fe0b1da28618dd17bda8b0a6bf5004554080 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-01T18:07:21+02:00 | GitHub | noreply@github.com | 2025-06-01T18:07:21+02:00 | | convert : fix nomic-bert-moe mask token (#13757) |
| 1525 | c496fe0b1da28618dd17bda8b0a6bf5004554080 | e57bb87cede38341963a7a884630dbfb09c7dc00 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-06-01T17:23:11+02:00 | GitHub | noreply@github.com | 2025-06-01T17:23:11+02:00 | | convert : fix vocab padding code for bert models (#13954) |
| 1526 | e57bb87cede38341963a7a884630dbfb09c7dc00 | f3a4b1659ccc94ebdf3590d472b518377bfecc7a | Aaron Teo | aaron.teo1@ibm.com | 2025-06-01T22:53:57+08:00 | GitHub | noreply@github.com | 2025-06-01T16:53:57+02:00 | | ggml: check if non-native endian model is being loaded (#13943) |
| 1527 | f3a4b1659ccc94ebdf3590d472b518377bfecc7a | 108009f5c7267b77292c11f788f602d6d7ae8fcd | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T12:23:14+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | sync : ggml |
| 1528 | 108009f5c7267b77292c11f788f602d6d7ae8fcd | d337252acf14a91a685c355fa4f3f599a8068207 | Kai Pastor | dg0yt@darc.de | 2025-05-31T12:49:55+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | vulkan : Remove unexpected ; (ggml/1253) |
| 1529 | d337252acf14a91a685c355fa4f3f599a8068207 | af6f91db470da543dc32bc00c057aad8f060dfdb | Kai Pastor | dg0yt@darc.de | 2025-05-31T12:39:19+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | cmake : Fix broken CMake error messages (ggml/1252) |
| 1530 | af6f91db470da543dc32bc00c057aad8f060dfdb | a7b8d35f780ab3e7639ae5345442fb1ad10f0163 | Radoslav Gerganov | rgerganov@gmail.com | 2025-05-30T09:11:09+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | ggml : remove ggml_graph_import and ggml_graph_export declarations (ggml/1247) |
| 1531 | a7b8d35f780ab3e7639ae5345442fb1ad10f0163 | 6eba72b71c677ce9aae33c422a6e67008137987a | Georgi Gerganov | ggerganov@gmail.com | 2025-05-29T13:29:50+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | sync : whisper.cpp (ggml/1250) |
| 1532 | 6eba72b71c677ce9aae33c422a6e67008137987a | fedf034a98378059274e4243d2a370a243335e73 | Radoslav Gerganov | rgerganov@gmail.com | 2025-05-29T08:34:46+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | ggml : install dynamic backends (ggml/1240) |
| 1533 | fedf034a98378059274e4243d2a370a243335e73 | 8726392d3d927fe3e504868d3548d694496ee37e | Daniel Tang | danielzgtg.opensource@gmail.com | 2025-05-27T20:58:46-04:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T13:43:57+03:00 | | ggml : Print backtrace on uncaught C++ exceptions (ggml/1232) |
| 1534 | 8726392d3d927fe3e504868d3548d694496ee37e | c04621711a893cbd09cff6c927cb005bc6749e36 | ddh0 | chemist-mulches-39@icloud.com | 2025-06-01T03:44:30-05:00 | GitHub | noreply@github.com | 2025-06-01T11:44:30+03:00 | | readme : update bindings (#13950) |
| 1535 | c04621711a893cbd09cff6c927cb005bc6749e36 | 0fc16b42e8793364918e830c234c1f3caec29dae | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T11:42:16+03:00 | GitHub | noreply@github.com | 2025-06-01T11:42:16+03:00 | | parallel : fix n_junk == 0 (#13952) |
| 1536 | 0fc16b42e8793364918e830c234c1f3caec29dae | 053b1539c02617eff744f89525ee57497c3c1fbe | Georgi Gerganov | ggerganov@gmail.com | 2025-06-01T11:39:27+03:00 | GitHub | noreply@github.com | 2025-06-01T11:39:27+03:00 | | kv-cache : split implementation in separate sources (#13920) |
| 1537 | 053b1539c02617eff744f89525ee57497c3c1fbe | b3a89c3d9e34c28c5be70d8b687a84775746d4a0 | Max Krasnyansky | quic_maxk@quicinc.com | 2025-05-31T15:39:19-07:00 | GitHub | noreply@github.com | 2025-05-31T15:39:19-07:00 | | threading: support for GGML_SCHED_PRIO_LOW, update thread info on Windows to avoid throttling (#12995) |
| 1538 | b3a89c3d9e34c28c5be70d8b687a84775746d4a0 | e15898d1c7e87f1b3d14e0a3c385eeb424b085fe | Jiří Podivín | 66251151+jpodivin@users.noreply.github.com | 2025-05-31T18:58:35+02:00 | GitHub | noreply@github.com | 2025-05-31T18:58:35+02:00 | | docs : Note about necessity of having libcurl installed for standard build. (#13945) |
| 1539 | e15898d1c7e87f1b3d14e0a3c385eeb424b085fe | 803f8baf4f741d2f0465c46c33f285886b97a071 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-31T08:26:10-07:00 | GitHub | noreply@github.com | 2025-05-31T08:26:10-07:00 | | server: allow unclosed thinking tags (#13931) |
| 1540 | 803f8baf4f741d2f0465c46c33f285886b97a071 | 3600cc2886956fc0a07ef6ad2f4128ccfdbc8c6f | Georgi Gerganov | ggerganov@gmail.com | 2025-05-31T15:58:33+03:00 | GitHub | noreply@github.com | 2025-05-31T15:58:33+03:00 | | llama : deprecate explicit kv_self defrag/update calls (#13921) |
| 1541 | 3600cc2886956fc0a07ef6ad2f4128ccfdbc8c6f | c7e0a2054b908c28bf93bb18d4b63ccbff2c4127 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-31T15:57:44+03:00 | GitHub | noreply@github.com | 2025-05-31T15:57:44+03:00 | | llama : use n_swa + n_ubatch cells for SWA cache (#13833) |
| 1542 | c7e0a2054b908c28bf93bb18d4b63ccbff2c4127 | 3f55f781f1da9241f209648636fa0426fe62a495 | igardev | 49397134+igardev@users.noreply.github.com | 2025-05-31T12:56:08+03:00 | GitHub | noreply@github.com | 2025-05-31T11:56:08+02:00 | | webui : Replace alert and confirm with custom modals. (#13711) |
| 1543 | 3f55f781f1da9241f209648636fa0426fe62a495 | 51fa76f172957c9916e43aeb581f82c533e12394 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-31T12:55:57+03:00 | GitHub | noreply@github.com | 2025-05-31T12:55:57+03:00 | | llama : auto-batch preparation (#13845) |
| 1544 | 51fa76f172957c9916e43aeb581f82c533e12394 | 12d0188c0dc6146ffde6d277a93f232ccbe699f8 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-31T10:14:29+02:00 | GitHub | noreply@github.com | 2025-05-31T10:14:29+02:00 | | mtmd : drop `_shared` from `libmtmd` name, merge helpers into libmtmd (⚠️ breaking change) (#13917) |
| 1545 | 12d0188c0dc6146ffde6d277a93f232ccbe699f8 | eb3949938e82a128855bc0676220bb2ce6e4228d | Georgi Gerganov | ggerganov@gmail.com | 2025-05-31T10:24:04+03:00 | GitHub | noreply@github.com | 2025-05-31T10:24:04+03:00 | | kv-cache : refactor + add llama_memory_state_i (#13746) |
| 1546 | eb3949938e82a128855bc0676220bb2ce6e4228d | e562eece7cb476276bfc4cbb18deb7c0369b2233 | Shawn yang | 137684499+Yangxiaoz@users.noreply.github.com | 2025-05-31T14:48:04+08:00 | GitHub | noreply@github.com | 2025-05-31T08:48:04+02:00 | | CUDA: add a prop in ggml_cuda_device_infor for distinguish iGPU or dGPU in cuda (#13856) (#13895) |
| 1547 | e562eece7cb476276bfc4cbb18deb7c0369b2233 | b47ab7b8e9c4f6154eb6abb6234866ab117d64c7 | Johannes Gäßler | johannesg@5d6.de | 2025-05-30T21:22:03+02:00 | GitHub | noreply@github.com | 2025-05-30T21:22:03+02:00 | | CUDA: fix typo in FlashAttention code (#13926) |
| 1548 | b47ab7b8e9c4f6154eb6abb6234866ab117d64c7 | dd665cc9d4b378626cde11ea0c4c91f514fdc994 | Diego Devesa | slarengh@gmail.com | 2025-05-30T09:56:19-07:00 | GitHub | noreply@github.com | 2025-05-30T18:56:19+02:00 | | sched : avoid changing cur_copy when a graph is already allocated (#13922) |
| 1549 | dd665cc9d4b378626cde11ea0c4c91f514fdc994 | df0c0c7d02f951dc639d48fd536f79200460ac83 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-30T19:38:07+03:00 | GitHub | noreply@github.com | 2025-05-30T19:38:07+03:00 | | parallel : increase the variability of the prompt lengths (#13927) |
| 1550 | df0c0c7d02f951dc639d48fd536f79200460ac83 | b49a8ff96b769b8a4c36d89fb783ec0135be582b | Diego Devesa | slarengh@gmail.com | 2025-05-30T07:37:18-07:00 | GitHub | noreply@github.com | 2025-05-30T16:37:18+02:00 | | cuda : prevent using split buffers with 3d/4d matrices (#13919) |
| 1551 | b49a8ff96b769b8a4c36d89fb783ec0135be582b | 53f925074de02c5304b00c14b4d6d8910c58667d | Akarshan Biswas | akarshan@menlo.ai | 2025-05-30T19:40:57+05:30 | GitHub | noreply@github.com | 2025-05-30T19:40:57+05:30 | | SYCL: Add mrope kernel (#13755) |
| 1552 | 53f925074de02c5304b00c14b4d6d8910c58667d | db38704f0133be7832123495fa8fc2601ea999d4 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-30T16:25:45+03:00 | GitHub | noreply@github.com | 2025-05-30T16:25:45+03:00 | | sync : vendor (#13901) |
| 1553 | db38704f0133be7832123495fa8fc2601ea999d4 | 07e4351ce663a7802b31f92fd7bc0e555c2044b6 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-30T14:50:43+02:00 | GitHub | noreply@github.com | 2025-05-30T14:50:43+02:00 | | convert : fix rwkv bos/eos token (#13844) |
| 1554 | 07e4351ce663a7802b31f92fd7bc0e555c2044b6 | 291f2b6913c7ef8350dbf0e77da38f7af131a08e | Xuan-Son Nguyen | son@huggingface.co | 2025-05-30T12:24:37+02:00 | GitHub | noreply@github.com | 2025-05-30T12:24:37+02:00 | | convert : allow partial update to the chkhsh pre-tokenizer list (#13847) |
| 1555 | 291f2b6913c7ef8350dbf0e77da38f7af131a08e | 2c90da4c7ec694797f524042aaafbb047a7e65ff | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-05-30T18:56:02+09:00 | GitHub | noreply@github.com | 2025-05-30T11:56:02+02:00 | | llama : add support for DistilBert (#13907) |
| 1556 | 2c90da4c7ec694797f524042aaafbb047a7e65ff | ec9e0301fef6476df83e94842c3b625501c95566 | zhangkaihuo | zhangkaihuo@gmail.com | 2025-05-30T16:31:48+08:00 | GitHub | noreply@github.com | 2025-05-30T10:31:48+02:00 | | llama : use llm_build_granite for minicpm (#13911) |
| 1557 | ec9e0301fef6476df83e94842c3b625501c95566 | e83ba3e460651b20a594e9f2f0f0bffb998d3ce1 | Christian Kastner | ckk@kvr.at | 2025-05-30T01:28:54+02:00 | GitHub | noreply@github.com | 2025-05-30T01:28:54+02:00 | | cmake: Guard GGML_CPU_ALL_VARIANTS by architecture (#13890) |
| 1558 | e83ba3e460651b20a594e9f2f0f0bffb998d3ce1 | 2b131621e60d8ec2cc961201beb6773ab37b6b69 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-29T21:42:31+02:00 | GitHub | noreply@github.com | 2025-05-29T21:42:31+02:00 | | llama : add support for jina-reranker-v2 (#13900) |
| 1559 | 2b131621e60d8ec2cc961201beb6773ab37b6b69 | 54a2c7a8cd8a32b44e3a98c2999b0f5c9114be5c | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-29T15:36:05+02:00 | GitHub | noreply@github.com | 2025-05-29T15:36:05+02:00 | | gguf-py : add support for sub_type (in arrays) in GGUFWriter add_key_value method (#13561) |
| 1560 | 54a2c7a8cd8a32b44e3a98c2999b0f5c9114be5c | 21fcc21ad5e5de2daa5da8d08dbbcc86b8d815d7 | Yibo Cai | yibo.cai@arm.com | 2025-05-29T19:39:20+08:00 | GitHub | noreply@github.com | 2025-05-29T14:39:20+03:00 | | arm64: optimize q4_k_q8_k kernel with i8mm (#13886) |
| 1561 | 21fcc21ad5e5de2daa5da8d08dbbcc86b8d815d7 | dd8ba93416f492c168f7e20f0b1434c08ebeaeb5 | Christian Kastner | ckk@kvr.at | 2025-05-29T12:50:25+02:00 | GitHub | noreply@github.com | 2025-05-29T12:50:25+02:00 | | cmake: Factor out CPU architecture detection (#13883) |
| 1562 | dd8ba93416f492c168f7e20f0b1434c08ebeaeb5 | 66c92061f5c34d2f9a630b7b50cca5ce3ebb4b16 | Vineel Abhinav | 131174187+vineelabhinav@users.noreply.github.com | 2025-05-29T14:48:43+05:30 | GitHub | noreply@github.com | 2025-05-29T12:18:43+03:00 | | ggml: aarch64: Implement SVE F32 kernels for Mamba Sequential Scan Algorithm (#13882) |
| 1563 | 66c92061f5c34d2f9a630b7b50cca5ce3ebb4b16 | 5ca82fc1d7f8cb27f3cbb3019e944a80d5219828 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-29T12:17:16+03:00 | GitHub | noreply@github.com | 2025-05-29T12:17:16+03:00 | | tests : remove json.hpp from a test (#13880) |
| 1564 | 5ca82fc1d7f8cb27f3cbb3019e944a80d5219828 | 6385b843a8dc8e15b8362196039720c58dd79fa2 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-29T10:00:57+02:00 | GitHub | noreply@github.com | 2025-05-29T10:00:57+02:00 | | convert : workaround for AutoConfig dummy labels (#13881) |
| 1565 | 6385b843a8dc8e15b8362196039720c58dd79fa2 | 1b8fb8152d0f3357a9c1d9b4c02f6cc5b9cb7232 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-29T08:15:01+02:00 | GitHub | noreply@github.com | 2025-05-29T08:15:01+02:00 | | llama : add RobertaForSequenceClassification reranker support (#13875) |
| 1566 | 1b8fb8152d0f3357a9c1d9b4c02f6cc5b9cb7232 | 53ae30640e131082d8d19bd80485b47c4553d551 | Vineel Abhinav | 131174187+vineelabhinav@users.noreply.github.com | 2025-05-29T11:31:33+05:30 | GitHub | noreply@github.com | 2025-05-29T09:01:33+03:00 | | ggml: aarch64: Implement SVE F32 kernels for vector functions (#13843) |
| 1567 | 53ae30640e131082d8d19bd80485b47c4553d551 | 763d06edb7dd5094ea58bad1d81e2e8d35033e34 | Beinsezii | 39478211+Beinsezii@users.noreply.github.com | 2025-05-28T14:50:20-07:00 | GitHub | noreply@github.com | 2025-05-28T23:50:20+02:00 | | gguf-py : fix SafetensorRemote return on undefined size (< 0) (#13841) |
| 1568 | 763d06edb7dd5094ea58bad1d81e2e8d35033e34 | 10961339b26bd2eff01d5479e8879f435da261b7 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-28T22:35:31+02:00 | GitHub | noreply@github.com | 2025-05-28T22:35:31+02:00 | | llama : fix KV shift for qwen2vl (#13870) |
| 1569 | 10961339b26bd2eff01d5479e8879f435da261b7 | d98f2a35fcf4a8d3e660ad48cd19e2a1f3d5b2ef | Xuan-Son Nguyen | son@huggingface.co | 2025-05-28T22:35:22+02:00 | GitHub | noreply@github.com | 2025-05-28T22:35:22+02:00 | | mtmd : move helpers to dedicated library (⚠️ breaking change) (#13866) |
| 1570 | d98f2a35fcf4a8d3e660ad48cd19e2a1f3d5b2ef | e0e3aa231d899720c2864d09cdb89d4c400eeb55 | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-05-28T15:46:47-03:00 | GitHub | noreply@github.com | 2025-05-28T15:46:47-03:00 | | ci: disable LLAMA_CURL for Linux cross-builds (#13871) |
| 1571 | e0e3aa231d899720c2864d09cdb89d4c400eeb55 | aa6dff05be25709bb218bf648951d690029c4b19 | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-05-29T02:01:58+09:00 | GitHub | noreply@github.com | 2025-05-28T19:01:58+02:00 | | llama : add support for BertForSequenceClassification reranker (#13858) |
| 1572 | aa6dff05be25709bb218bf648951d690029c4b19 | c962ae3382a1e759c8517a229549ee53685313a1 | Đinh Trọng Huy | 77562200+huydt84@users.noreply.github.com | 2025-05-28T23:34:18+09:00 | GitHub | noreply@github.com | 2025-05-28T16:34:18+02:00 | | convert: small addition to support LlamaModel (#13838) |
| 1573 | c962ae3382a1e759c8517a229549ee53685313a1 | a3938fb53d0b5183db4f5a6db21fd9122a0e4780 | Sky | Iflyinskyin2013@gmail.com | 2025-05-28T22:33:54+08:00 | GitHub | noreply@github.com | 2025-05-28T16:33:54+02:00 | | server: fix remove 'image_url'/'input_audio' json-object effectlly for 'llama_params' in multimodal-model-mode (#13853) |
| 1574 | a3938fb53d0b5183db4f5a6db21fd9122a0e4780 | f7873fc698c09047e2873630ab7e7730a0bfb224 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-28T16:12:35+02:00 | GitHub | noreply@github.com | 2025-05-28T16:12:35+02:00 | | convert : fix qwen omni conversion (#13859) |
| 1575 | f7873fc698c09047e2873630ab7e7730a0bfb224 | a68247439bd6fb756cc93ad2817e55a02aa0b100 | Alex Fanthome | xfanth@gmail.com | 2025-05-28T14:49:28+01:00 | GitHub | noreply@github.com | 2025-05-28T15:49:28+02:00 | | tests : change umlaut test (#11600) |
| 1576 | a68247439bd6fb756cc93ad2817e55a02aa0b100 | 26b79b6cb3e7840ff15729350e95907e19f9f480 | Johannes Gäßler | johannesg@5d6.de | 2025-05-28T13:33:37+02:00 | GitHub | noreply@github.com | 2025-05-28T13:33:37+02:00 | | CUDA: fix FA tg at long context for CC >= 8.9 (#13852) |
| 1577 | 26b79b6cb3e7840ff15729350e95907e19f9f480 | 1e8659e65ad0f640674d2785d02b556ab938728c | Xuan-Son Nguyen | son@huggingface.co | 2025-05-28T10:05:54+02:00 | GitHub | noreply@github.com | 2025-05-28T10:05:54+02:00 | | convert : fix tensor naming conflict for llama 4 vision (#13836) |
| 1578 | 1e8659e65ad0f640674d2785d02b556ab938728c | a3c30846e410c91c11d7bf80978795a03bb03dee | leo-pony | nengjunma@outlook.com | 2025-05-28T11:54:20+08:00 | GitHub | noreply@github.com | 2025-05-28T11:54:20+08:00 | | CANN: Add SOC TYPE printing in cmake configuration (#13837) |
| 1579 | a3c30846e410c91c11d7bf80978795a03bb03dee | 1701d4c54f93c0d203bf92ea7202a94b037c2338 | lhez | quic_lih@quicinc.com | 2025-05-27T12:56:08-07:00 | GitHub | noreply@github.com | 2025-05-27T12:56:08-07:00 | | opencl: add new ops - `argsort`, `div`, `sub`, `addrows`, `sigmoid`, `group_norm` (#13787) |
| 1580 | 1701d4c54f93c0d203bf92ea7202a94b037c2338 | bef817638780dcbd8c0c80c97b7a4f8e92c8fe74 | lhez | quic_lih@quicinc.com | 2025-05-27T12:53:14-07:00 | GitHub | noreply@github.com | 2025-05-27T12:53:14-07:00 | | opencl: mark `mul_mat` `f32f32` as supporting non-contiguous tensors (#13790) |
| 1581 | bef817638780dcbd8c0c80c97b7a4f8e92c8fe74 | 34b7c0439ed0f98575cc4689dfecd98991dee8be | Jeff Bolz | jbolz@nvidia.com | 2025-05-27T11:39:07-05:00 | GitHub | noreply@github.com | 2025-05-27T18:39:07+02:00 | | vulkan: use timestamp queries for GGML_VULKAN_PERF (#13817) |
| 1582 | 34b7c0439ed0f98575cc4689dfecd98991dee8be | f3101a8cc665f73217c752a10a7042889275cfbc | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T19:08:44+03:00 | GitHub | noreply@github.com | 2025-05-27T19:08:44+03:00 | | cmake : add llama-cparams.cpp to build (#13832) |
| 1583 | f3101a8cc665f73217c752a10a7042889275cfbc | 1c49c70d07ef87635daa5e8fdd0b5bfd88493dd3 | Akarshan Biswas | akarshan@menlo.ai | 2025-05-27T20:52:59+05:30 | GitHub | noreply@github.com | 2025-05-27T20:52:59+05:30 | | SYCL: add gelu_erf kernel (#13749) |
| 1584 | 1c49c70d07ef87635daa5e8fdd0b5bfd88493dd3 | a8ea03d8ad9c984fe6cfafead183ab188f8cbeb0 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T18:04:38+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T18:05:33+03:00 | | sync : ggml |
| 1585 | a8ea03d8ad9c984fe6cfafead183ab188f8cbeb0 | 05f6ac6283a7859a03cd59405504262c85c4bf06 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-27T15:53:55+02:00 | GitHub | noreply@github.com | 2025-05-27T15:53:55+02:00 | | ggml : add ggml_repeat_4d (#13824) |
| 1586 | 05f6ac6283a7859a03cd59405504262c85c4bf06 | bc583e3c63c04a11d287c108ea9e6a515ead0423 | xctan | xc-tan@outlook.com | 2025-05-27T21:21:36+08:00 | GitHub | noreply@github.com | 2025-05-27T16:21:36+03:00 | | ggml : riscv: add xtheadvector support (#13720) |
| 1587 | bc583e3c63c04a11d287c108ea9e6a515ead0423 | 72b090da2c50e540143fd312a2f9aa5f151e6136 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-27T14:06:10+02:00 | GitHub | noreply@github.com | 2025-05-27T14:06:10+02:00 | | mtmd : support Qwen 2.5 Omni (input audio+vision, no audio output) (#13784) |
| 1588 | 72b090da2c50e540143fd312a2f9aa5f151e6136 | 7fe03e7446ed4388539d3bcc013ae4e851b8d1e2 | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-05-27T08:52:40-03:00 | GitHub | noreply@github.com | 2025-05-27T08:52:40-03:00 | | docs: remove link for llama-cli function calling (#13810) |
| 1589 | 7fe03e7446ed4388539d3bcc013ae4e851b8d1e2 | 952f3953c1b61cc70e79e536c42ddce6a5ea5ea7 | Christian Kastner | ckk@kvr.at | 2025-05-27T13:18:39+02:00 | GitHub | noreply@github.com | 2025-05-27T13:18:39+02:00 | | ggml-cpu: x86 feature detection is specific to x86 (#13811) |
| 1590 | 952f3953c1b61cc70e79e536c42ddce6a5ea5ea7 | 81713121ee56755ba4a147d1a1e3fa248ec31a66 | Diego Devesa | slarengh@gmail.com | 2025-05-27T04:05:18-07:00 | GitHub | noreply@github.com | 2025-05-27T13:05:18+02:00 | | ggml : allow CUDA graphs when using pipeline parallelism (#13814) |
| 1591 | 81713121ee56755ba4a147d1a1e3fa248ec31a66 | f9cd68398baf2ba8af4725ca9ed00bef205e6706 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T13:49:41+03:00 | GitHub | noreply@github.com | 2025-05-27T13:49:41+03:00 | | kv-cells : track min/max used cells and per-sequence positions (#13808) |
| 1592 | f9cd68398baf2ba8af4725ca9ed00bef205e6706 | 4f81b33e324b1669b282eb81104e4a1131be7dce | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T12:07:52+03:00 | GitHub | noreply@github.com | 2025-05-27T12:07:52+03:00 | | sampling : make sure samplers return at least 1 token (#13822) |
| 1593 | 4f81b33e324b1669b282eb81104e4a1131be7dce | cdf94a18023c92f41808ec874ba577d914674717 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-27T09:40:59+03:00 | GitHub | noreply@github.com | 2025-05-27T09:40:59+03:00 | | llama : validate seq id batch input (#13809) |
| 1594 | cdf94a18023c92f41808ec874ba577d914674717 | a26c4cc11ec7c6574e3691e90ecdbd67deeea35b | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-26T14:34:27-07:00 | GitHub | noreply@github.com | 2025-05-26T22:34:27+01:00 | | server: --offline mode (#13804) |
| 1595 | a26c4cc11ec7c6574e3691e90ecdbd67deeea35b | 4265a87b59ebfc25f35adbf4db3b608995b0a78a | Georgi Gerganov | ggerganov@gmail.com | 2025-05-26T22:24:01+03:00 | GitHub | noreply@github.com | 2025-05-26T22:24:01+03:00 | | scripts : add option to compare commits in Debug (#13806) |
| 1596 | 4265a87b59ebfc25f35adbf4db3b608995b0a78a | 6f180b915c9ed9ec0c240b5dcd64644988fb5e82 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-26T22:14:52+03:00 | GitHub | noreply@github.com | 2025-05-26T22:14:52+03:00 | | cuda : avoid cuGetErrorString (#13791) |
| 1597 | 6f180b915c9ed9ec0c240b5dcd64644988fb5e82 | 03f582ae8fccecff225c30a2802461b44761e822 | Akarshan Biswas | akarshan@menlo.ai | 2025-05-26T21:10:36+05:30 | GitHub | noreply@github.com | 2025-05-26T21:10:36+05:30 | | SYCL: Add non contiguous support in RMS_NORM and NORM kernels (#13611) |
| 1598 | 03f582ae8fccecff225c30a2802461b44761e822 | 88c125f2acce0e25e5fc8481ab0681415fc64a10 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-26T08:03:57-07:00 | GitHub | noreply@github.com | 2025-05-26T16:03:57+01:00 | | server: fix streaming crashes (#13786) |
| 1599 | 88c125f2acce0e25e5fc8481ab0681415fc64a10 | d74e94c1b3ac02db01e47008e1f8d371029d81ac | standby24x7 | standby24x7@gmail.com | 2025-05-26T23:55:24+09:00 | GitHub | noreply@github.com | 2025-05-26T16:55:24+02:00 | | examples/training: Fix file name in README (#13803) |
| 1600 | d74e94c1b3ac02db01e47008e1f8d371029d81ac | f13847cfb560d92492b480cca2c5d6aa9473cde3 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-26T06:56:49-07:00 | GitHub | noreply@github.com | 2025-05-26T14:56:49+01:00 | | `server`: fix format of streamed tool call deltas (diff name, fix id location) (#13800) |
| 1601 | f13847cfb560d92492b480cca2c5d6aa9473cde3 | 79c137f77677b3c8ee3c60a7da033721b938399a | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-26T06:16:37-07:00 | GitHub | noreply@github.com | 2025-05-26T14:16:37+01:00 | | server: fix regression on streamed non-chat completion w/ stops (#13785) |
| 1602 | 79c137f77677b3c8ee3c60a7da033721b938399a | 22229314fc46b2f741bb21b12cde71f6c6a60b52 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-26T14:03:54+03:00 | GitHub | noreply@github.com | 2025-05-26T14:03:54+03:00 | | examples : allow extracting embeddings from decoder contexts (#13797) |
| 1603 | 22229314fc46b2f741bb21b12cde71f6c6a60b52 | 9012eb9b454a82eaa4cd77ae904c0ea391e4db42 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-26T12:57:50+03:00 | GitHub | noreply@github.com | 2025-05-26T12:57:50+03:00 | | llama : clarify deprecation message (#13794) |
| 1604 | 9012eb9b454a82eaa4cd77ae904c0ea391e4db42 | fef693dc6b959a8e8ba11558fbeaad0b264dd457 | Romain Biessy | romain.biessy@codeplay.com | 2025-05-26T10:28:53+02:00 | GitHub | noreply@github.com | 2025-05-26T10:28:53+02:00 | | sycl: Add more debug prints (#13640) |
| 1605 | fef693dc6b959a8e8ba11558fbeaad0b264dd457 | 2d38b6e4004fb1c341723e657fb1e71a4d3fb473 | Jeff Bolz | jbolz@nvidia.com | 2025-05-25T23:02:07-05:00 | GitHub | noreply@github.com | 2025-05-26T06:02:07+02:00 | | vulkan: mark IM2COL as supporting non-contig (#13783) |
| 1606 | 2d38b6e4004fb1c341723e657fb1e71a4d3fb473 | e121edc4324a640be11b7e567edd39b721b0f8e4 | Bizhao Shi | 37729561+shibizhao@users.noreply.github.com | 2025-05-26T10:20:18+08:00 | GitHub | noreply@github.com | 2025-05-26T10:20:18+08:00 | | CANN: Add the basic supports of Flash Attention kernel (#13627) |
| 1607 | e121edc4324a640be11b7e567edd39b721b0f8e4 | 2f099b510f460374acd52742b494595e3e3442d3 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-26T00:30:51+01:00 | GitHub | noreply@github.com | 2025-05-26T00:30:51+01:00 | | `server`: add `--reasoning-budget 0` to disable thinking (incl. qwen3 w/ enable_thinking:false) (#13771) |
| 1608 | 2f099b510f460374acd52742b494595e3e3442d3 | aa50ba462f63759e1731227f20f5009cfc2a0f16 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-25T19:02:18+02:00 | GitHub | noreply@github.com | 2025-05-25T18:02:18+01:00 | | webui : bump max upload file size to 500MB (#13779) |
| 1609 | aa50ba462f63759e1731227f20f5009cfc2a0f16 | de2ef53a4b2c0d703749a309d19fe68fd8f1b9ac | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-25T16:22:29+02:00 | GitHub | noreply@github.com | 2025-05-25T16:22:29+02:00 | | tests : improve UGM tokenizer test coverage (#13773) |
| 1610 | de2ef53a4b2c0d703749a309d19fe68fd8f1b9ac | c508256db2de2b032e19c8ed833f4683c827c9a1 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-25T16:34:36+03:00 | GitHub | noreply@github.com | 2025-05-25T16:34:36+03:00 | | kv-cache : rework kv_cell (#13706) |
| 1611 | c508256db2de2b032e19c8ed833f4683c827c9a1 | 40aaa8a403df5dcfb515eb2fd7a817fd99779526 | Percy Piper | piper.percy@googlemail.com | 2025-05-25T13:35:53+01:00 | GitHub | noreply@github.com | 2025-05-25T15:35:53+03:00 | | rpc : Fix build on OpenBSD (#13541) |
| 1612 | 40aaa8a403df5dcfb515eb2fd7a817fd99779526 | a08c1d2845dc279d58d826ef3c3ecad97cbbcef7 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-25T14:06:32+02:00 | GitHub | noreply@github.com | 2025-05-25T14:06:32+02:00 | | mtmd : add support for Qwen2-Audio and SeaLLM-Audio (#13760) |
| 1613 | a08c1d2845dc279d58d826ef3c3ecad97cbbcef7 | d785f9c1fd1a1929c6d0e2a0b12cae5db867908b | ddpasa | 112642920+ddpasa@users.noreply.github.com | 2025-05-25T14:04:49+02:00 | GitHub | noreply@github.com | 2025-05-25T14:04:49+02:00 | | docs : add Moondream2 pre-quantized link (#13745) |
| 1614 | d785f9c1fd1a1929c6d0e2a0b12cae5db867908b | 4032ca406632212b075e3c61c2a4476128321410 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-25T10:45:49+01:00 | GitHub | noreply@github.com | 2025-05-25T10:45:49+01:00 | | server: fix/test add_generation_prompt (#13770) |
| 1615 | 4032ca406632212b075e3c61c2a4476128321410 | 515fdbf7ed839dfe6a24aeb6225936609a7f6d6d | Piotr Jasiukajtis | estibi@me.com | 2025-05-25T10:29:43+02:00 | GitHub | noreply@github.com | 2025-05-25T10:29:43+02:00 | | llama : add support for Qwen3 MoE tied word embeddings (#13768) |
| 1616 | f5cd27b71da3ac375a04a41643d14fc779a8057b | a2d02d5793fd9af7a7224773456501691b95fd02 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-25T01:48:08+01:00 | GitHub | noreply@github.com | 2025-05-25T01:48:08+01:00 | | `server`: streaming of tool calls and thoughts when `--jinja` is on (#12379) |
| 1617 | a2d02d5793fd9af7a7224773456501691b95fd02 | 17fc817b58942569c43aa63299cedde217b61f68 | Diego Devesa | slarengh@gmail.com | 2025-05-24T15:55:16-07:00 | GitHub | noreply@github.com | 2025-05-25T00:55:16+02:00 | | releases : bundle llvm omp library in windows release (#13763) |
| 1618 | 17fc817b58942569c43aa63299cedde217b61f68 | 2bd1b30f6979235ec67b95c183b9b77baa7ab9ce | Diego Devesa | slarengh@gmail.com | 2025-05-24T13:27:03-07:00 | GitHub | noreply@github.com | 2025-05-24T22:27:03+02:00 | | releases : enable openmp in windows cpu backend build (#13756) |
| 1619 | 2bd1b30f6979235ec67b95c183b9b77baa7ab9ce | 259469c4b57c1a32606353bcac52ba683424a990 | Diego Devesa | slarengh@gmail.com | 2025-05-24T13:26:47-07:00 | GitHub | noreply@github.com | 2025-05-24T22:26:47+02:00 | | ggml-cpu : set openmp wait time if not set (#13758) |
| 1620 | 259469c4b57c1a32606353bcac52ba683424a990 | 4c32832c59e97d96c19349fdc92763e8949fda83 | 0cc4m | picard12@live.de | 2025-05-24T16:49:12+02:00 | GitHub | noreply@github.com | 2025-05-24T16:49:12+02:00 | | Move GLM4 f32 attention fix to the correct function (#13750) |
| 1621 | 4c32832c59e97d96c19349fdc92763e8949fda83 | c3a2624339187e89c4f65fd72a5fe7103968b5ad | Xuan-Son Nguyen | son@huggingface.co | 2025-05-24T13:06:47+02:00 | GitHub | noreply@github.com | 2025-05-24T13:06:47+02:00 | | ggml : add ggml_gelu_erf() CUDA kernel (#13719) |
| 1622 | c3a2624339187e89c4f65fd72a5fe7103968b5ad | ffd0eae60b7642e942c8245b89642527e36099e6 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-24T12:29:09+02:00 | GitHub | noreply@github.com | 2025-05-24T12:29:09+02:00 | | vocab : fix ugm tokenizer precision (#13743) |
| 1623 | ffd0eae60b7642e942c8245b89642527e36099e6 | b775345d788ac16260e7eef49e11fe57ee5677f7 | Johannes Gäßler | johannesg@5d6.de | 2025-05-24T11:46:19+02:00 | GitHub | noreply@github.com | 2025-05-24T11:46:19+02:00 | | CUDA: fix race condition in FA vector kernels (#13742) |
| 1624 | b775345d788ac16260e7eef49e11fe57ee5677f7 | a70a8a69c2ad3de8d5525ad2b185b098011be2e8 | Diego Devesa | slarengh@gmail.com | 2025-05-23T13:14:00-07:00 | GitHub | noreply@github.com | 2025-05-23T23:14:00+03:00 | | ci : enable winget package updates (#13734) |
| 1625 | a70a8a69c2ad3de8d5525ad2b185b098011be2e8 | d13d0f6135803822ec1cd7e3efb49360b88a1bdf | Diego Devesa | slarengh@gmail.com | 2025-05-23T13:09:38-07:00 | GitHub | noreply@github.com | 2025-05-23T22:09:38+02:00 | | ci : add winget package updater (#13732) |
| 1626 | d13d0f6135803822ec1cd7e3efb49360b88a1bdf | 8a2afb7520bbc8f9fa1bbe314d5f2807eb0116b2 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-23T20:16:13+03:00 | GitHub | noreply@github.com | 2025-05-23T20:16:13+03:00 | | hparams : initialize arrays (#13728) |
| 1627 | 8a2afb7520bbc8f9fa1bbe314d5f2807eb0116b2 | 9ecf3e66a38f20e41d221f62f05b17f70d040839 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-23T17:07:04+02:00 | GitHub | noreply@github.com | 2025-05-23T17:07:04+02:00 | | llama : allow custom list of swa_layers (#13726) |
| 1628 | 9ecf3e66a38f20e41d221f62f05b17f70d040839 | faaaff5f947064d13ef8b98659d81a1384c3e57b | Xuan-Son Nguyen | son@huggingface.co | 2025-05-23T11:03:47+02:00 | GitHub | noreply@github.com | 2025-05-23T11:03:47+02:00 | | server : support audio input (#13714) |
| 1629 | faaaff5f947064d13ef8b98659d81a1384c3e57b | e16c4731c7a6bfa8f61b5b39c77245a37e7fc232 | Chenguang Li | 757486878@qq.com | 2025-05-23T16:47:53+08:00 | GitHub | noreply@github.com | 2025-05-23T16:47:53+08:00 | | CANN: Support MUL_MAT_ID for q8_0 and q4_0 (#13705) |
| 1630 | e16c4731c7a6bfa8f61b5b39c77245a37e7fc232 | 1dcd01960c33f52afb782b3d850fc2149a08cc6b | Xuan-Son Nguyen | son@huggingface.co | 2025-05-23T08:12:48+02:00 | GitHub | noreply@github.com | 2025-05-23T08:12:48+02:00 | | ggml : fix the order of ggml_unary_op (#13718) |
| 1631 | 1dcd01960c33f52afb782b3d850fc2149a08cc6b | c10ed6cbcc3904b7d6bbcd137589eebd899aa38f | Jeff Bolz | jbolz@nvidia.com | 2025-05-23T00:45:02-04:00 | GitHub | noreply@github.com | 2025-05-23T06:45:02+02:00 | | vulkan: support CPY from any type to itself (#13695) |
| 1632 | c10ed6cbcc3904b7d6bbcd137589eebd899aa38f | a127ff17800452075bb323277e3b9f8db8612ced | Jeff Bolz | jbolz@nvidia.com | 2025-05-23T00:33:45-04:00 | GitHub | noreply@github.com | 2025-05-23T06:33:45+02:00 | | vulkan: Disable coopmat/coopmat2/bfloat extensions if glslc doesn't support it (#13696) |
| 1633 | a127ff17800452075bb323277e3b9f8db8612ced | 3079e9ac8e04ef6eddeb0c164d72edb6b6fd2df5 | Judd | 4046440+foldl@users.noreply.github.com | 2025-05-23T12:33:08+08:00 | GitHub | noreply@github.com | 2025-05-23T06:33:08+02:00 | | use LOG_WARN to replace `std::cerr` (#13657) |
| 1634 | 3079e9ac8e04ef6eddeb0c164d72edb6b6fd2df5 | 8a1d206f1d2b4e45918b589f3165b4be232f7ba8 | Diego Devesa | slarengh@gmail.com | 2025-05-22T15:21:37-07:00 | GitHub | noreply@github.com | 2025-05-23T00:21:37+02:00 | | release : fix windows hip release (#13707) |
| 1635 | 8a1d206f1d2b4e45918b589f3165b4be232f7ba8 | 797990c4bca0dca5be295c63e3fb2800dc0a69c2 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-22T22:21:07+03:00 | GitHub | noreply@github.com | 2025-05-22T22:21:07+03:00 | | tts : fix n_ubatch + make WavTokenizer cache-less (#13713) |
| 1636 | 797990c4bca0dca5be295c63e3fb2800dc0a69c2 | ab86335760ebb441574eb47f886fa1ee302e2131 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-22T20:42:48+02:00 | GitHub | noreply@github.com | 2025-05-22T20:42:48+02:00 | | mtmd : add ultravox audio input (#13623) |
| 1637 | ab86335760ebb441574eb47f886fa1ee302e2131 | cc74d5be990e37f201591fd868a92e64abdbf902 | Aaron Teo | aaron.teo1@ibm.com | 2025-05-23T02:31:29+08:00 | GitHub | noreply@github.com | 2025-05-22T21:31:29+03:00 | | common: Include torch package for s390x (#13699) |
| 1638 | cc74d5be990e37f201591fd868a92e64abdbf902 | 5be24af73d77ea23d399726f1d2a01a70ee86331 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-22T16:33:39+03:00 | GitHub | noreply@github.com | 2025-05-22T16:33:39+03:00 | | server : pad small embedding batches (#13692) |
| 1639 | 5be24af73d77ea23d399726f1d2a01a70ee86331 | d394a9aedc50a13b7f6373416f7c1ccabfe79c32 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-22T14:25:05+02:00 | GitHub | noreply@github.com | 2025-05-22T14:25:05+02:00 | | gguf-py : correct charsmap parameter typing (#13701) |
| 1640 | d394a9aedc50a13b7f6373416f7c1ccabfe79c32 | 6b56a64690a318fcabcd7739ac7e314d44785ea8 | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-05-22T13:54:43+02:00 | GitHub | noreply@github.com | 2025-05-22T12:54:43+01:00 | | sycl : Remove waits from function calls (#13702) |
| 1641 | 6b56a64690a318fcabcd7739ac7e314d44785ea8 | a4e8912dfd4604be1e39bc86ba4c0b02969967ef | Ewan Crawford | ewan@codeplay.com | 2025-05-22T09:24:09+01:00 | GitHub | noreply@github.com | 2025-05-22T16:24:09+08:00 | | SYCL: Avoid using with SYCL-Graph for unsupported nodes (#13587) |
| 1642 | a4e8912dfd4604be1e39bc86ba4c0b02969967ef | edbf42edfdabb9cea72ae12137570cf48f5d8076 | Henry Linjamäki | henry.mikael.linjamaki@intel.com | 2025-05-22T02:21:45+03:00 | GitHub | noreply@github.com | 2025-05-21T16:21:45-07:00 | | opencl: Add support for multiple devices (#12622) |
| 1643 | edbf42edfdabb9cea72ae12137570cf48f5d8076 | d643bb2c798df9c2cd61067d2692b1cd417df402 | Henry Linjamäki | henry.mikael.linjamaki@intel.com | 2025-05-21T23:21:17+03:00 | GitHub | noreply@github.com | 2025-05-21T13:21:17-07:00 | | opencl: fix couple crashes (#12795) |
| 1644 | d643bb2c798df9c2cd61067d2692b1cd417df402 | 8e186ef0e764c7a620e402d1f76ebad60bf31c49 | Diego Devesa | slarengh@gmail.com | 2025-05-21T13:09:57-07:00 | GitHub | noreply@github.com | 2025-05-21T22:09:57+02:00 | | releases : build CPU backend separately (windows) (#13642) |
| 1645 | 8e186ef0e764c7a620e402d1f76ebad60bf31c49 | 5fbfe384d4659f81c47a477eb8ee97692c7ffef9 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-21T20:00:49+03:00 | GitHub | noreply@github.com | 2025-05-21T20:00:49+03:00 | | hparams : support models for which all layers use SWA (#13682) |
| 1646 | 5fbfe384d4659f81c47a477eb8ee97692c7ffef9 | c76532e7ba128bb097bf6836bf0f5592e1b56b76 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-21T19:46:56+03:00 | GitHub | noreply@github.com | 2025-05-21T19:46:56+03:00 | | server : improve error reporting (#13680) |
| 1647 | c76532e7ba128bb097bf6836bf0f5592e1b56b76 | 2aa777d86d3f7bb80b93e226f1c25e47825f6a83 | antichristHater | 142441588+antichristHater@users.noreply.github.com | 2025-05-21T19:40:35+03:00 | GitHub | noreply@github.com | 2025-05-21T18:40:35+02:00 | | convert : add qwen2vl support for unsloth merges (#13686) |
| 1648 | 2aa777d86d3f7bb80b93e226f1c25e47825f6a83 | eb0f5c28d37126baa756117d5bdaadc62e03344e | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-21T16:57:38+02:00 | GitHub | noreply@github.com | 2025-05-21T16:57:38+02:00 | | examples : switch retrieval to llama_encode (#13685) |
| 1649 | eb0f5c28d37126baa756117d5bdaadc62e03344e | cf4cb59e64d72a1b4c781f71a74de5756a4e2376 | Emmanuel Ferdman | emmanuelferdman@gmail.com | 2025-05-21T17:33:54+03:00 | GitHub | noreply@github.com | 2025-05-21T16:33:54+02:00 | | gguf-py : display the invalid gguf type (#13687) |
| 1650 | cf4cb59e64d72a1b4c781f71a74de5756a4e2376 | 0d5c74216170ef97e5e7511563837263f2d1a496 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-21T16:26:33+02:00 | GitHub | noreply@github.com | 2025-05-21T16:26:33+02:00 | | ggml : add ggml_gelu_erf() (#13667) |
| 1651 | 0d5c74216170ef97e5e7511563837263f2d1a496 | 42158ae2e8ead667a83f07247321ce85f32ace66 | Robin Davidsson | 40024429+R-Dson@users.noreply.github.com | 2025-05-21T15:15:27+02:00 | GitHub | noreply@github.com | 2025-05-21T15:15:27+02:00 | | server : Add the endpoints /api/tags and /api/chat (#13659) |
| 1652 | 42158ae2e8ead667a83f07247321ce85f32ace66 | 797f2ac0625b22edeff03cc30e0f988da6b6b068 | Dorin-Andrei Geman | doringeman@gmail.com | 2025-05-21T16:07:57+03:00 | GitHub | noreply@github.com | 2025-05-21T15:07:57+02:00 | | server : fix first message identification (#13634) |
| 1653 | 797f2ac0625b22edeff03cc30e0f988da6b6b068 | b44890df2e4fad0ece1d5366dcbc8bedae23b658 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-21T15:11:13+03:00 | GitHub | noreply@github.com | 2025-05-21T15:11:13+03:00 | | kv-cache : simplify the interface (#13660) |
| 1654 | b44890df2e4fad0ece1d5366dcbc8bedae23b658 | 33983057d0f578aca74ba15eccc3de9c267a5ff6 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-21T13:09:21+03:00 | GitHub | noreply@github.com | 2025-05-21T13:09:21+03:00 | | model : disable SWA for Phi models (#13676) |
| 1655 | 33983057d0f578aca74ba15eccc3de9c267a5ff6 | fb1cab201c6c4cd36731f100df74eece2f4706fa | R0CKSTAR | yeahdongcn@gmail.com | 2025-05-21T09:58:49+08:00 | GitHub | noreply@github.com | 2025-05-21T09:58:49+08:00 | | musa: Upgrade MUSA SDK version to rc4.0.1 and use mudnn::Unary::IDENTITY op to accelerate D2D memory copy (#13647) |
| 1656 | fb1cab201c6c4cd36731f100df74eece2f4706fa | b7a17463ec190aeee7b9077c606c910fb4688b84 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-05-20T21:35:16Z | GitHub | noreply@github.com | 2025-05-20T21:35:16Z | | vulkan: fix warnings (#13626) |
| 1657 | b7a17463ec190aeee7b9077c606c910fb4688b84 | be0239693c1530a18496086331fc18d8a9adbad1 | l3utterfly | gc.pthzfoldr@gmail.com | 2025-05-21T00:55:30+08:00 | GitHub | noreply@github.com | 2025-05-20T18:55:30+02:00 | | mtmd-helper : bug fix to token batching in mtmd (#13650) |
| 1658 | be0239693c1530a18496086331fc18d8a9adbad1 | a4090d1174aed22dde5cacce2a4c27656b987a2f | Georgi Gerganov | ggerganov@gmail.com | 2025-05-20T19:21:04+03:00 | GitHub | noreply@github.com | 2025-05-20T19:21:04+03:00 | | model : fix llama4 graph (#13663) |
| 1659 | a4090d1174aed22dde5cacce2a4c27656b987a2f | b69f1647f9953cb3773266d2c83a92fd0e7e6d66 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-20T16:13:16+03:00 | GitHub | noreply@github.com | 2025-05-20T16:13:16+03:00 | | llama : remove llama_kv_cache_view API + remove deprecated (#13653) |
| 1660 | b69f1647f9953cb3773266d2c83a92fd0e7e6d66 | 759e37b0d89bc4bd1bce860dc5f3c3052e08575c | Johannes Gäßler | johannesg@5d6.de | 2025-05-20T14:45:07+02:00 | GitHub | noreply@github.com | 2025-05-20T14:45:07+02:00 | | CUDA: skip fully masked-out KV in FA vec kernel (#13584) |
| 1661 | 759e37b0d89bc4bd1bce860dc5f3c3052e08575c | 4245e622e0cc3af1ca3056104e465dc4d303bd7d | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-20T12:03:17+02:00 | GitHub | noreply@github.com | 2025-05-20T12:03:17+02:00 | | tests : avoid github urls due to throttling (#13654) |
| 1662 | 4245e622e0cc3af1ca3056104e465dc4d303bd7d | c9c64dee572c5eedf073a6c6eb5d92d8283f7639 | Svetlozar Georgiev | 55534064+sgeor255@users.noreply.github.com | 2025-05-20T10:34:15+01:00 | GitHub | noreply@github.com | 2025-05-20T11:34:15+02:00 | | sycl: disable reorder for sycl mulmat (#13536) |
| 1663 | c9c64dee572c5eedf073a6c6eb5d92d8283f7639 | c00a2634bec49805c6b31438b1a6006a2bd793cb | 0cc4m | picard12@live.de | 2025-05-20T10:11:56+02:00 | GitHub | noreply@github.com | 2025-05-20T10:11:56+02:00 | | Set GLM4 blk.*.attn_output.weight, kqv_out-* matmul to GGML_PREC_F32 to fix infinity values in output (#13639) |
| 1664 | c00a2634bec49805c6b31438b1a6006a2bd793cb | e298d2fbd082a52c0f6ed02729f94e9bf630cf17 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-20T10:41:40+03:00 | GitHub | noreply@github.com | 2025-05-20T10:41:40+03:00 | | metal : fix typo in FA kernel comments (#13651) |
| 1665 | e298d2fbd082a52c0f6ed02729f94e9bf630cf17 | f0adb80bf7c2c0d80abb04f4533b5513622d9964 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-20T08:05:46+03:00 | GitHub | noreply@github.com | 2025-05-20T08:05:46+03:00 | | kv-cache : add SWA support (#13194) |
| 1666 | f0adb80bf7c2c0d80abb04f4533b5513622d9964 | f7c9429c85748cde9599499601ba48d0057722e6 | Xinpeng Dou | 15529241576@163.com | 2025-05-20T11:43:43+08:00 | GitHub | noreply@github.com | 2025-05-20T11:43:43+08:00 | | CANN: Update CANN model support (#13162) |
| 1667 | f7c9429c85748cde9599499601ba48d0057722e6 | 1dfbf2cf3a9f15193dd893396d07762bbd2c4785 | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-05-20T02:54:43+02:00 | GitHub | noreply@github.com | 2025-05-20T08:54:43+08:00 | | sycl : Overcoming workaround for mmap() allocation on Windows (#13482) |
| 1668 | 1dfbf2cf3a9f15193dd893396d07762bbd2c4785 | 8960efd0a65e5a4830a14f6937c10703534f2569 | psocolovsky | 50770545+psocolovsky@users.noreply.github.com | 2025-05-19T21:17:36+02:00 | GitHub | noreply@github.com | 2025-05-19T21:17:36+02:00 | | common : add load_progress_callback (#13617) |
| 1669 | 8960efd0a65e5a4830a14f6937c10703534f2569 | 725f23f1f3f0d3adf49f95d8dfa6e7c74adff149 | 0cc4m | picard12@live.de | 2025-05-19T17:54:08+02:00 | GitHub | noreply@github.com | 2025-05-19T17:54:08+02:00 | | Vulkan: Add f32 accumulator support to quantized mul mat to fix GLM4 32B incoherence (#13607) |
| 1670 | 725f23f1f3f0d3adf49f95d8dfa6e7c74adff149 | 92ecdcc06a4c405a415bcaa0cb772bc560aa23b1 | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-05-19T14:38:20+01:00 | GitHub | noreply@github.com | 2025-05-19T14:38:20+01:00 | | sycl : backend documentation review (#13544) |
| 1671 | 92ecdcc06a4c405a415bcaa0cb772bc560aa23b1 | f71f40a2847d4c9f57b86cd206e0a27b2bfb6d1c | Xuan-Son Nguyen | son@huggingface.co | 2025-05-19T13:04:14+02:00 | GitHub | noreply@github.com | 2025-05-19T13:04:14+02:00 | | mtmd : add vision support for llama 4 (#13282) |
| 1672 | f71f40a2847d4c9f57b86cd206e0a27b2bfb6d1c | d30cb5a7fa17362c47e94a023276f169916e0d03 | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-05-19T11:46:09+01:00 | GitHub | noreply@github.com | 2025-05-19T11:46:09+01:00 | | ci : upgraded oneAPI version in SYCL workflows and dockerfile (#13532) |
| 1673 | d30cb5a7fa17362c47e94a023276f169916e0d03 | 6c35981a643d4f74a286268b62f65bfa156af2e6 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-19T12:50:29+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-19T13:29:56+03:00 | | sync : ggml |
| 1674 | 6c35981a643d4f74a286268b62f65bfa156af2e6 | 8b5e19aea6ce9fe4452598663924373234041440 | Johannes Gäßler | johannesg@5d6.de | 2025-05-19T09:33:35+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-19T13:29:56+03:00 | | mnist: fix segmentation fault (ggml/1227) |
| 1675 | 8b5e19aea6ce9fe4452598663924373234041440 | 60aea028b575ab54e0dfbb31db641bb93d2ae8c4 | Diego Devesa | slarengh@gmail.com | 2025-05-18T18:30:13-07:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-19T13:29:56+03:00 | | ggml : fix apple OS check in ggml_print_backtrace (ggml/1229) |
| 1676 | 60aea028b575ab54e0dfbb31db641bb93d2ae8c4 | 9c55e5c5c24990dd17fd6c8f4f2159052d2b06f1 | Daniel Tang | danielzgtg.opensource@gmail.com | 2025-05-17T19:06:26-04:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-19T13:29:56+03:00 | | ggml : Fix missing backtrace on Linux (ggml/1228) |
| 1677 | 9c55e5c5c24990dd17fd6c8f4f2159052d2b06f1 | 33d7aed4a83e7bbecac4af535208d4ed3e1ca8fd | Nick | 0x0b4ac@gmail.com | 2025-05-19T18:25:41+08:00 | GitHub | noreply@github.com | 2025-05-19T13:25:41+03:00 | | fix: check model pointer validity before use (#13631) |
| 1678 | 33d7aed4a83e7bbecac4af535208d4ed3e1ca8fd | 6a2bc8bfb7cd502e5ebc72e36c97a6f848c21c2c | Chenguang Li | 757486878@qq.com | 2025-05-19T14:21:17+08:00 | GitHub | noreply@github.com | 2025-05-19T14:21:17+08:00 | | CANN: Support MOE Model MUL_MAT_ID (#13042) |
| 1679 | 6a2bc8bfb7cd502e5ebc72e36c97a6f848c21c2c | e3a7cf6c5bf6a0a24217f88607b06e4405a2b5d9 | Isaac McFadyen | isaac@imcf.me | 2025-05-17T17:59:48-04:00 | GitHub | noreply@github.com | 2025-05-17T23:59:48+02:00 | | server : added --no-prefill-assistant flag (#13608) |
| 1680 | e3a7cf6c5bf6a0a24217f88607b06e4405a2b5d9 | 518329b2d4434d0683c30b741b4f4cf5734b5f99 | Gilad S. | 7817232+giladgd@users.noreply.github.com | 2025-05-17T21:26:43+03:00 | GitHub | noreply@github.com | 2025-05-17T15:26:43-03:00 | | cmake: use the current build config for vulkan-shaders-gen (#13595) |
| 1681 | 518329b2d4434d0683c30b741b4f4cf5734b5f99 | 2f5a4e1e09567e09be8b84a5f976334de9539d17 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-17T12:58:55+03:00 | GitHub | noreply@github.com | 2025-05-17T12:58:55+03:00 | | parallel : add option for non-shared and larger prompts (#13598) |
| 1682 | 2f5a4e1e09567e09be8b84a5f976334de9539d17 | 4f41ee11d6a4ddb02e4644922bf0b56880ee6f55 | Jeff Bolz | jbolz@nvidia.com | 2025-05-17T16:14:55+09:00 | GitHub | noreply@github.com | 2025-05-17T09:14:55+02:00 | | vulkan: move common FA code to flash_attn_base.comp (#13556) |
| 1683 | 4f41ee11d6a4ddb02e4644922bf0b56880ee6f55 | 3e0be1cacef290c99cbb99ceaa433b4344f87355 | Jeff Bolz | jbolz@nvidia.com | 2025-05-17T15:35:47+09:00 | GitHub | noreply@github.com | 2025-05-17T08:35:47+02:00 | | vulkan: use scalar FA rather than coopmat2 when N==1 (#13554) |
| 1684 | 3e0be1cacef290c99cbb99ceaa433b4344f87355 | 6aa892ec2aa7fe0c93e87c4b970d83a942fb9454 | Z | coffeevampirebusiness@gmail.com | 2025-05-16T14:56:28-06:00 | GitHub | noreply@github.com | 2025-05-16T22:56:28+02:00 | | llguidance : official v0.7.20 release (no actual changes) [noci] (#13594) |
| 1685 | 6aa892ec2aa7fe0c93e87c4b970d83a942fb9454 | aea9f8b4e73876739bf88fb705502f291294e469 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-16T21:50:00+02:00 | GitHub | noreply@github.com | 2025-05-16T21:50:00+02:00 | | server : do not return error out of context (with ctx shift disabled) (#13577) |
| 1686 | aea9f8b4e73876739bf88fb705502f291294e469 | 06c1e4abc1e50ef6dcf6369f9685ca8fe82d21fe | Xuan-Son Nguyen | son@huggingface.co | 2025-05-16T21:49:01+02:00 | GitHub | noreply@github.com | 2025-05-16T21:49:01+02:00 | | webui : improve accessibility for visually impaired people (#13551) |
| 1687 | 06c1e4abc1e50ef6dcf6369f9685ca8fe82d21fe | 415e40a357002d5d9b7a97ef20ccea9f4ae04c80 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-16T20:04:18+02:00 | GitHub | noreply@github.com | 2025-05-16T20:04:18+02:00 | | readme : add list of dependencies and their license (#13591) |
| 1688 | 415e40a357002d5d9b7a97ef20ccea9f4ae04c80 | 654a67794f7a420c6931de2f4c145eaaade5dd16 | Diego Devesa | slarengh@gmail.com | 2025-05-16T10:36:51-07:00 | GitHub | noreply@github.com | 2025-05-16T19:36:51+02:00 | | releases : use arm version of curl for arm releases (#13592) |
| 1689 | 654a67794f7a420c6931de2f4c145eaaade5dd16 | 5364ae4ba53cc6367b8c8bf78876839122ca4e57 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-16T20:32:58+03:00 | GitHub | noreply@github.com | 2025-05-16T20:32:58+03:00 | | metal : add FA-vec kernel for head size 64 (#13583) |
| 1690 | 5364ae4ba53cc6367b8c8bf78876839122ca4e57 | 7c07ac244d59c833bf209582f6df019a77cdda59 | Diego Devesa | slarengh@gmail.com | 2025-05-16T07:38:07-07:00 | GitHub | noreply@github.com | 2025-05-16T16:38:07+02:00 | | llama : print hint when loading a model when no backends are loaded (#13589) |
| 1691 | 7c07ac244d59c833bf209582f6df019a77cdda59 | 0a338ed013c23aecdce6449af736a35a465fa60f | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-16T14:54:23+02:00 | GitHub | noreply@github.com | 2025-05-16T14:54:23+02:00 | | ci : add ppc64el to build-linux-cross (#13575) |
| 1692 | 0a338ed013c23aecdce6449af736a35a465fa60f | bc098c3cf0aac93e57e9bda9d95e2456acf88894 | Łukasz Ślusarczyk | 112692748+lslusarczyk@users.noreply.github.com | 2025-05-16T12:15:29+02:00 | GitHub | noreply@github.com | 2025-05-16T18:15:29+08:00 | | sycl : fixed compilation warnings (#13582) |
| 1693 | bc098c3cf0aac93e57e9bda9d95e2456acf88894 | c6a2c9e7411f54b0770b319740561bbd6a2ebd27 | Olivier Chafik | olivier.chafik@gmail.com | 2025-05-15T23:29:10+01:00 | GitHub | noreply@github.com | 2025-05-15T23:29:10+01:00 | | minja: sync (qwen3) (#13573) |
| 1694 | c6a2c9e7411f54b0770b319740561bbd6a2ebd27 | 07ad2b6db3c117868728e2cdfe3b711473d66e14 | Diego Devesa | slarengh@gmail.com | 2025-05-15T10:13:11-07:00 | GitHub | noreply@github.com | 2025-05-15T19:13:11+02:00 | | gguf : use ggml log system (#13571) |
| 1695 | 07ad2b6db3c117868728e2cdfe3b711473d66e14 | c531edfa34fec8074d0b424396cdd450df23b612 | Daniel Tang | danielzgtg.opensource@gmail.com | 2025-05-15T12:47:10-04:00 | GitHub | noreply@github.com | 2025-05-15T18:47:10+02:00 | | gguf-py : fix disconnect-before-connect in editor-gui (#13569) |
| 1696 | c531edfa34fec8074d0b424396cdd450df23b612 | 02cdd2d8b092b5a4bb18e013c6887ce49ba20ac5 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-15T17:40:07+02:00 | GitHub | noreply@github.com | 2025-05-15T17:40:07+02:00 | | convert : fix conversion for llama 4 (#13567) |
| 1697 | 02cdd2d8b092b5a4bb18e013c6887ce49ba20ac5 | 64bb51cf90d3eede8c150a23d59a0c718b78065b | Atharva Dubey | atharva.dubey@codeplay.com | 2025-05-15T16:39:52+01:00 | GitHub | noreply@github.com | 2025-05-15T17:39:52+02:00 | | sycl: simplify bin_bcast_kernel (#13383) |
| 1698 | 64bb51cf90d3eede8c150a23d59a0c718b78065b | 9c404ed54c3c8d8d2aa3153313766c8286739387 | Svetlozar Georgiev | 55534064+sgeor255@users.noreply.github.com | 2025-05-15T16:35:44+01:00 | GitHub | noreply@github.com | 2025-05-15T17:35:44+02:00 | | sycl: reordered Q4_K MMVQ (#13109) |
| 1699 | 9c404ed54c3c8d8d2aa3153313766c8286739387 | 6c8b91500e75df6664278d1e9af3e39e8a2fb0d0 | Łukasz Ślusarczyk | 112692748+lslusarczyk@users.noreply.github.com | 2025-05-15T16:53:41+02:00 | GitHub | noreply@github.com | 2025-05-15T16:53:41+02:00 | | sycl: use oneDNN for matrices multiplication (#12972) |
| 1700 | 6c8b91500e75df6664278d1e9af3e39e8a2fb0d0 | 3cc1f1f1d24472a6558c942b1c78989ff4b0e569 | Diego Devesa | slarengh@gmail.com | 2025-05-15T06:46:55-07:00 | GitHub | noreply@github.com | 2025-05-15T15:46:55+02:00 | | llama-bench : fix -ot with dl backends (#13563) |
| 1701 | 3cc1f1f1d24472a6558c942b1c78989ff4b0e569 | c753d7bed0dc2cc2d5c42dfa9806cba91748392e | Xuan-Son Nguyen | son@huggingface.co | 2025-05-15T14:24:50+02:00 | GitHub | noreply@github.com | 2025-05-15T14:24:50+02:00 | | webui : handle PDF input (as text or image) + convert pasted long content to file (#13562) |
| 1702 | c753d7bed0dc2cc2d5c42dfa9806cba91748392e | b2838049ccf50859cea4c390e81405c2fb01e820 | Piotr Wilkin (ilintar) | piotr.wilkin@syndatis.com | 2025-05-15T08:40:58+02:00 | GitHub | noreply@github.com | 2025-05-15T08:40:58+02:00 | | server : proper error handling for missing elements in messages array (OpenAI compatible backend) (#13540) |
| 1703 | b2838049ccf50859cea4c390e81405c2fb01e820 | aa48e373f256df9608395ee6881e7010840d6202 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-15T05:57:02+03:00 | GitHub | noreply@github.com | 2025-05-15T05:57:02+03:00 | | bench : handle decode errors (#13548) |
| 1704 | aa48e373f256df9608395ee6881e7010840d6202 | e3a9421b78da5c810d812e6348a405f5acc39f34 | Olivier Chafik | ochafik@users.noreply.github.com | 2025-05-15T02:39:51+01:00 | GitHub | noreply@github.com | 2025-05-15T02:39:51+01:00 | | `server`: inject date_string in llama 3.x template + fix date for firefunction v2 (#12802) |
| 1705 | e3a9421b78da5c810d812e6348a405f5acc39f34 | 5ab5d5fb256aa23f8ab4de8463515ea627f43006 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-14T23:15:15+03:00 | GitHub | noreply@github.com | 2025-05-14T23:15:15+03:00 | | kv-cache : fix out-of-bounds view during reserve graph (#13547) |
| 1706 | 5ab5d5fb256aa23f8ab4de8463515ea627f43006 | 3198405e98530c683d12fd123e3920f2bd2aafa5 | Yibo Cai | cyb70289@gmail.com | 2025-05-15T03:53:52+08:00 | GitHub | noreply@github.com | 2025-05-14T21:53:52+02:00 | | arm64: optimize q6_k_q8_k kernel with i8mm (#13519) |
| 1707 | 3198405e98530c683d12fd123e3920f2bd2aafa5 | f5170c1d7a66222ca7c75d2022fec3ed87257e0b | Olivier Chafik | ochafik@users.noreply.github.com | 2025-05-14T19:50:57+01:00 | GitHub | noreply@github.com | 2025-05-14T19:50:57+01:00 | | `common`: add partial regex support (#12808) |
| 1708 | f5170c1d7a66222ca7c75d2022fec3ed87257e0b | 017f10b5fa630a013ec4f9936e410a60d4f460d5 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-14T20:22:49+02:00 | GitHub | noreply@github.com | 2025-05-14T21:22:49+03:00 | | editorconfig : fix trailing whitespace from #13542 (#13546) |
| 1709 | 017f10b5fa630a013ec4f9936e410a60d4f460d5 | 4696d5674999dc10a7fb8c27b33406a929f7463a | Gilad S. | 7817232+giladgd@users.noreply.github.com | 2025-05-14T19:18:18+03:00 | GitHub | noreply@github.com | 2025-05-14T19:18:18+03:00 | | fix: crash when calling `llama_state_get_size` on a context without a KV cache (#13542) |
| 1710 | 4696d5674999dc10a7fb8c27b33406a929f7463a | b7d26720821823e23e2273a99e38398d511242e9 | Johannes Gäßler | johannesg@5d6.de | 2025-05-14T16:41:02+02:00 | GitHub | noreply@github.com | 2025-05-14T16:41:02+02:00 | | CUDA: fix crash on large batch size for quant. MoE (#13537) |
| 1711 | b7d26720821823e23e2273a99e38398d511242e9 | 6da34fa27620fa56e3334172e023f4f2533df51f | Diego Devesa | slarengh@gmail.com | 2025-05-14T07:12:36-07:00 | GitHub | noreply@github.com | 2025-05-14T16:12:36+02:00 | | llama : fix quantize with dl backends (#13539) |
| 1712 | 6da34fa27620fa56e3334172e023f4f2533df51f | 5e7d95e22e386d316f7f659b74c9c34b65507912 | Johannes Gäßler | johannesg@5d6.de | 2025-05-14T16:08:20+02:00 | GitHub | noreply@github.com | 2025-05-14T16:08:20+02:00 | | CUDA: faster Deepseek FA, add Turing support (#13435) |
| 1713 | 5e7d95e22e386d316f7f659b74c9c34b65507912 | 053174436f3b3cbb01077f26ff17809537b832c7 | Gabe Goodhart | ghart@us.ibm.com | 2025-05-14T06:53:59-06:00 | GitHub | noreply@github.com | 2025-05-14T15:53:59+03:00 | | fix: Move build_inp_pos to the top of the graph section for build_granite (#13538) |
| 1714 | 053174436f3b3cbb01077f26ff17809537b832c7 | 360a9c98e13d35f322b4c5b1309aab0cc90ed82b | Georgi Gerganov | ggerganov@gmail.com | 2025-05-14T15:42:10+03:00 | GitHub | noreply@github.com | 2025-05-14T15:42:10+03:00 | | server : passthrough the /models endpoint during loading (#13535) |
| 1715 | 360a9c98e13d35f322b4c5b1309aab0cc90ed82b | 09d13d94fb169093095416ac6323551c37b75339 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-14T13:35:07+02:00 | GitHub | noreply@github.com | 2025-05-14T13:35:07+02:00 | | server : fix cache_tokens bug with no cache_prompt (#13533) |
| 1716 | 09d13d94fb169093095416ac6323551c37b75339 | 24e86cae7219b0f3ede1d5abdf5bf3ad515cccb8 | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-05-14T07:53:57-03:00 | GitHub | noreply@github.com | 2025-05-14T07:53:57-03:00 | | cmake: simplify vulkan shader test logic (#13263) |
| 1717 | 24e86cae7219b0f3ede1d5abdf5bf3ad515cccb8 | bb1681fbd532eba26ae4c14cd8be884c8afeb31c | Jeff Bolz | jbolz@nvidia.com | 2025-05-14T18:55:26+09:00 | GitHub | noreply@github.com | 2025-05-14T11:55:26+02:00 | | vulkan: KHR_coopmat flash attention (#13506) |
| 1718 | bb1681fbd532eba26ae4c14cd8be884c8afeb31c | d486dd3e8ecfba0170f8632a1530540de560f4a3 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-14T10:26:12+02:00 | GitHub | noreply@github.com | 2025-05-14T10:26:12+02:00 | | webui : use fflate for more deterministic gzip compress (#13525) |
| 1719 | d486dd3e8ecfba0170f8632a1530540de560f4a3 | 21ca987fba504d273ea28ebc4d3e5b3736a11c8e | Luca Stefani | luca.stefani.ge1@gmail.com | 2025-05-14T10:07:31+02:00 | GitHub | noreply@github.com | 2025-05-14T10:07:31+02:00 | | webui: Allow pasting file from clipboard (#13526) |
| 1720 | 21ca987fba504d273ea28ebc4d3e5b3736a11c8e | be1d4a13db26750fac702ceb3af88ae4f39dc9f4 | ddpasa | 112642920+ddpasa@users.noreply.github.com | 2025-05-14T09:59:12+02:00 | GitHub | noreply@github.com | 2025-05-14T09:59:12+02:00 | | docs: Update link to ggml-org in multimodal.md (#13513) |
| 1721 | be1d4a13db26750fac702ceb3af88ae4f39dc9f4 | ab3971f2a0a526564bfde0e4cc8a5b90f9d33ad2 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-14T08:41:01+02:00 | GitHub | noreply@github.com | 2025-05-14T08:41:01+02:00 | | scripts : fix compare-llama-bench.py show parameter (#13514) |
| 1722 | ab3971f2a0a526564bfde0e4cc8a5b90f9d33ad2 | e5c834f718a32b7584f142799bbf508fddb9021c | Jeff Bolz | jbolz@nvidia.com | 2025-05-14T13:15:50+09:00 | GitHub | noreply@github.com | 2025-05-14T06:15:50+02:00 | | vulkan: workaround FA compile failures on macos (#13517) |
| 1723 | e5c834f718a32b7584f142799bbf508fddb9021c | 71bdbdb58757d508557e6d8b387f666cdfb25c5e | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-05-13T18:12:31+01:00 | GitHub | noreply@github.com | 2025-05-13T19:12:31+02:00 | | quantize : improve tensor-type pattern matching (#13033) |
| 1724 | 71bdbdb58757d508557e6d8b387f666cdfb25c5e | f0995d28ce3d15095b6845d94ce4465e46575873 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-13T17:07:21+02:00 | GitHub | noreply@github.com | 2025-05-13T17:07:21+02:00 | | clip : clip.h become private API (⚠️ breaking change) (#13510) |
| 1725 | f0995d28ce3d15095b6845d94ce4465e46575873 | c252e0c4097b34666e5a81db9d0450d71fa3098f | Georgi Gerganov | ggerganov@gmail.com | 2025-05-13T18:04:39+03:00 | GitHub | noreply@github.com | 2025-05-13T18:04:39+03:00 | | metal : use FA-vec kernel up to batch size 20 (#13496) |
| 1726 | c252e0c4097b34666e5a81db9d0450d71fa3098f | 4f711afed5e7ef4304b567c8888ee1aa60e868eb | Georgi Gerganov | ggerganov@gmail.com | 2025-05-13T18:04:00+03:00 | GitHub | noreply@github.com | 2025-05-13T18:04:00+03:00 | | metal : optimize multi-sequence FA vec kernel (#13493) |
| 1727 | 4f711afed5e7ef4304b567c8888ee1aa60e868eb | b89d605a91dee0a518ecd582f8991a07d523e2fa | Dan Johansson | dan.johansson@arm.com | 2025-05-13T17:02:28+02:00 | GitHub | noreply@github.com | 2025-05-13T18:02:28+03:00 | | ggml-cpu: Update KleidiAI to v1.6 and fix include directives (#13509) |
| 1728 | b89d605a91dee0a518ecd582f8991a07d523e2fa | b4726345aca49e2ad62d615e6e370b3dbad6434f | Georgi Gerganov | ggerganov@gmail.com | 2025-05-13T18:01:53+03:00 | GitHub | noreply@github.com | 2025-05-13T18:01:53+03:00 | | batched-bench : fix pp batch contents (#13492) |
| 1729 | b4726345aca49e2ad62d615e6e370b3dbad6434f | bf7937112058f2815fc3825a9ff7b536ecafa3bb | Xuan-Son Nguyen | son@huggingface.co | 2025-05-13T15:33:58+02:00 | GitHub | noreply@github.com | 2025-05-13T15:33:58+02:00 | | mtmd : remove libllava, remove clip-quantize-cli (⚠️ breaking change) (#13460) |
| 1730 | bf7937112058f2815fc3825a9ff7b536ecafa3bb | d590cd4c244e5f260c42c290b83a358b9d86d763 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-13T15:31:12+02:00 | GitHub | noreply@github.com | 2025-05-13T15:31:12+02:00 | | scripts : support arbitrary input file formats in compare-llama-bench.py (#13455) |
| 1731 | d590cd4c244e5f260c42c290b83a358b9d86d763 | 1e2809bc4b5d8db2a9ed12ac872eca832c53f5fd | Gabe Goodhart | ghart@us.ibm.com | 2025-05-13T07:12:01-06:00 | GitHub | noreply@github.com | 2025-05-13T15:12:01+02:00 | | model : Granite MoE shared (#13269) |
| 1732 | 1e2809bc4b5d8db2a9ed12ac872eca832c53f5fd | cf0a43bb6490bd49344775abb22ba26f8047cb54 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-13T14:01:45+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-13T14:02:28+03:00 | | sync : ggml |
| 1733 | cf0a43bb6490bd49344775abb22ba26f8047cb54 | f0d46ef15717cd609a7b69cf6190edde64d466c8 | Diego Devesa | slarengh@gmail.com | 2025-05-12T15:31:37-07:00 | GitHub | noreply@github.com | 2025-05-13T00:31:37+02:00 | | llama-bench : add defrag-thold, check for invalid ranges (#13487) |
| 1734 | f0d46ef15717cd609a7b69cf6190edde64d466c8 | de4c07f93783a1a96456a44dc16b9db538ee1618 | lhez | quic_lih@quicinc.com | 2025-05-12T13:13:49-07:00 | GitHub | noreply@github.com | 2025-05-12T13:13:49-07:00 | | opencl: remove unnecessary assert for `add` (#13257) |
| 1735 | de4c07f93783a1a96456a44dc16b9db538ee1618 | 10d2af0eaa0aafd7c6577b279dfa5221ff44a63f | Xuan-Son Nguyen | son@huggingface.co | 2025-05-12T15:06:51+02:00 | GitHub | noreply@github.com | 2025-05-12T15:06:51+02:00 | | clip : cap max image size 1024 for qwen vl model (#13478) |
| 1736 | 10d2af0eaa0aafd7c6577b279dfa5221ff44a63f | 064cc596ac44308dc326a17c9e3163c34a6f29d1 | Johannes Gäßler | johannesg@5d6.de | 2025-05-12T14:44:49+02:00 | GitHub | noreply@github.com | 2025-05-12T14:44:49+02:00 | | llama/ggml: add LLM training support (#10544) |
| 1737 | 064cc596ac44308dc326a17c9e3163c34a6f29d1 | 91159ee9df84edb72ccf874e353d2414f3a8c60d | Georgi Gerganov | ggerganov@gmail.com | 2025-05-12T15:12:27+03:00 | GitHub | noreply@github.com | 2025-05-12T15:12:27+03:00 | | context : fix state io for memory-less contexts (#13470) |
| 1738 | 91159ee9df84edb72ccf874e353d2414f3a8c60d | 22cdab343b63edd0906f1c132616746085f81983 | Anudit Nagar | nagaranudit@gmail.com | 2025-05-12T18:56:42+07:00 | GitHub | noreply@github.com | 2025-05-12T13:56:42+02:00 | | server : allow content to be null in oaicompat_completion_params_parse (#13477) |
| 1739 | 22cdab343b63edd0906f1c132616746085f81983 | a71a4075cdb8e81eaaa834e48ffda88b038286bc | Diego Devesa | slarengh@gmail.com | 2025-05-12T13:08:22+02:00 | GitHub | noreply@github.com | 2025-05-12T13:08:22+02:00 | | llama-bench : accept ranges for integer parameters (#13410) |
| 1740 | a71a4075cdb8e81eaaa834e48ffda88b038286bc | 95e18884fc7ea4031f70f1a518d5d1df616e5717 | Dan Johansson | dan.johansson@arm.com | 2025-05-12T13:06:19+02:00 | GitHub | noreply@github.com | 2025-05-12T13:06:19+02:00 | | ggml-cpu: Integrate fp32=bf16xbf16 SME KleidiAI kernel (#13053) |
| 1741 | 95e18884fc7ea4031f70f1a518d5d1df616e5717 | df8491922f0ea4e032338186350dff2fb6b2b4e3 | Johannes Gäßler | johannesg@5d6.de | 2025-05-12T10:51:21+02:00 | GitHub | noreply@github.com | 2025-05-12T10:51:21+02:00 | | CUDA: fix misaligned synchronization in FA (#13469) |
| 1742 | df8491922f0ea4e032338186350dff2fb6b2b4e3 | 14492144c286bbf38bb1903128403d9e2ebad54c | Xuan-Son Nguyen | son@huggingface.co | 2025-05-12T10:29:13+02:00 | GitHub | noreply@github.com | 2025-05-12T10:29:13+02:00 | | ggml : add mrope kernel for metal (#13457) |
| 1743 | 14492144c286bbf38bb1903128403d9e2ebad54c | c104023994d36a8e791fc6a43789b84fd552cefc | Atharva Dubey | atharva.dubey@codeplay.com | 2025-05-12T06:15:32+01:00 | GitHub | noreply@github.com | 2025-05-12T13:15:32+08:00 | | enable dpcpp nightly builds with libraries (#13406) |
| 1744 | c104023994d36a8e791fc6a43789b84fd552cefc | 9a390c4829cd3058d26a2e2c09d16e3fd12bf1b1 | City | 125218114+city96@users.noreply.github.com | 2025-05-12T00:39:06+02:00 | GitHub | noreply@github.com | 2025-05-12T00:39:06+02:00 | | mtmd : Use RMS norm for InternVL 3 38B and 78B mmproj (#13459) |
| 1745 | 9a390c4829cd3058d26a2e2c09d16e3fd12bf1b1 | 09232370fc6426aa5dd9be01a8271b9c28f5af3a | Anthony Umfer | aumfer@gmail.com | 2025-05-11T11:08:26-04:00 | GitHub | noreply@github.com | 2025-05-11T17:08:26+02:00 | | tools : fix uninitialized llama_batch in server (#13436) |
| 1746 | 09232370fc6426aa5dd9be01a8271b9c28f5af3a | 7474e00b34629e9cd8b06bc87ad935584ea30f8e | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-11T16:20:39+02:00 | GitHub | noreply@github.com | 2025-05-11T16:20:39+02:00 | | scripts : exit compare-llama-bench.py gracefully when there's nothing to compare (#13451) |
| 1747 | 7474e00b34629e9cd8b06bc87ad935584ea30f8e | 7f323a589f8684c0eb722e7309074cb5eac0c8b5 | Johannes Gäßler | johannesg@5d6.de | 2025-05-11T16:09:33+02:00 | GitHub | noreply@github.com | 2025-05-11T16:09:33+02:00 | | CUDA: fix crash with partial offloading of MoE (#13439) |
| 1748 | 7f323a589f8684c0eb722e7309074cb5eac0c8b5 | 3eac209319a6726fd9687c6188fc6b916b65953d | David Huang | 1969802+hjc4869@users.noreply.github.com | 2025-05-11T20:18:39+08:00 | GitHub | noreply@github.com | 2025-05-11T14:18:39+02:00 | | Add `--no-op-offload` to improve `-ot` pp perf in MoE models like llama4 400B (#13386) |
| 1749 | 3eac209319a6726fd9687c6188fc6b916b65953d | a634d75d1bb4034b8c945042ceea268b5b895730 | City | 125218114+city96@users.noreply.github.com | 2025-05-11T11:35:52+02:00 | GitHub | noreply@github.com | 2025-05-11T11:35:52+02:00 | | mtmd : support InternVL 3 38B and 78B mmproj (#13443) |
| 1750 | a634d75d1bb4034b8c945042ceea268b5b895730 | 62d4250e52917dc4d634f054b1e2183ed7dee944 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-11T11:34:23+02:00 | GitHub | noreply@github.com | 2025-05-11T11:34:23+02:00 | | mtmd : move helpers to dedicated file (#13442) |
| 1751 | 62d4250e52917dc4d634f054b1e2183ed7dee944 | 0208355f42bdab88a08507ead4a6302790a08323 | Thomas Germer | 99991@users.noreply.github.com | 2025-05-10T22:26:46+02:00 | GitHub | noreply@github.com | 2025-05-10T22:26:46+02:00 | | docs : Fix typo in InternVL3 model name (#13440) |
| 1752 | 0208355f42bdab88a08507ead4a6302790a08323 | d2a4ef05c60506ee48e7375eb36f2257de7ab0d2 | Johannes Gäßler | johannesg@5d6.de | 2025-05-10T22:22:48+02:00 | GitHub | noreply@github.com | 2025-05-10T22:22:48+02:00 | | CUDA: fix race conditions FlashAttention kernels (#13438) |
| 1753 | d2a4ef05c60506ee48e7375eb36f2257de7ab0d2 | 15e6125a397f6086c1dfdf7584acdb7c730313dc | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-10T22:08:07+02:00 | GitHub | noreply@github.com | 2025-05-10T22:08:07+02:00 | | vocab : add ByteDance-Seed/Seed-Coder (#13423) |
| 1754 | 15e6125a397f6086c1dfdf7584acdb7c730313dc | 3b24d26c22b9d92b5aa39930988700c84e524cf4 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-10T19:57:54+02:00 | GitHub | noreply@github.com | 2025-05-10T19:57:54+02:00 | | mtmd : add hard limit on image resolution for qwen2vl / qwen2.5vl (#13434) |
| 1755 | 3b24d26c22b9d92b5aa39930988700c84e524cf4 | 43dfd741a5da77982e0a94d1e522a2481dfb914d | Xuan-Son Nguyen | son@huggingface.co | 2025-05-10T18:44:49+02:00 | GitHub | noreply@github.com | 2025-05-10T18:44:49+02:00 | | server : update docs (#13432) |
| 1756 | 43dfd741a5da77982e0a94d1e522a2481dfb914d | b064a51a4e7246fa43a3307f58e59f65e6be170a | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-10T17:19:52+02:00 | GitHub | noreply@github.com | 2025-05-10T17:19:52+02:00 | | llguidance : set tokenizer slices to default (#13424) |
| 1757 | b064a51a4e7246fa43a3307f58e59f65e6be170a | 053367d149f778cdabc356ee3024494e0dd53223 | Thammachart Chinvarapon | 1731496+Thammachart@users.noreply.github.com | 2025-05-10T21:34:48+07:00 | GitHub | noreply@github.com | 2025-05-10T16:34:48+02:00 | | ci: free_disk_space flag enabled for intel variant (#13426) |
| 1758 | 053367d149f778cdabc356ee3024494e0dd53223 | d8919424f1dee7dc1638349c616f2ef5d2ee16fb | Xuan-Son Nguyen | son@huggingface.co | 2025-05-10T16:26:42+02:00 | GitHub | noreply@github.com | 2025-05-10T16:26:42+02:00 | | mtmd : support InternVL 2.5 and 3 (#13422) |
| 1759 | d8919424f1dee7dc1638349c616f2ef5d2ee16fb | 7fef11766cdeb9fa7bbbe3db13580616b7d3d599 | Johannes Gäßler | johannesg@5d6.de | 2025-05-10T09:16:52+02:00 | GitHub | noreply@github.com | 2025-05-10T09:16:52+02:00 | | CUDA: fix FlashAttention on Turing (#13415) |
| 1760 | 7fef11766cdeb9fa7bbbe3db13580616b7d3d599 | dc1d2adfc0f4de84da7923866d00781ec5c4e666 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-10T08:16:29+02:00 | GitHub | noreply@github.com | 2025-05-10T08:16:29+02:00 | | arg : add env var to control mmproj (#13416) |
| 1761 | dc1d2adfc0f4de84da7923866d00781ec5c4e666 | 7c28a74e0783f4bb74a246fb9f19bf212139e365 | Jeff Bolz | jbolz@nvidia.com | 2025-05-09T23:07:07-07:00 | GitHub | noreply@github.com | 2025-05-10T08:07:07+02:00 | | vulkan: scalar flash attention implementation (#13324) |
| 1762 | 7c28a74e0783f4bb74a246fb9f19bf212139e365 | 33eff4024084d1f0c8441b79f7208a52fad79858 | Helton Reis | 47722840+HRKings@users.noreply.github.com | 2025-05-09T17:15:39-03:00 | GitHub | noreply@github.com | 2025-05-09T23:15:39+03:00 | | chore(llguidance): use tagged version that does not break the build (#13413) |
| 1763 | 33eff4024084d1f0c8441b79f7208a52fad79858 | 17512a94d636c4b6c1332370acb3e5af3ca70918 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-09T19:29:37+02:00 | GitHub | noreply@github.com | 2025-05-09T19:29:37+02:00 | | server : vision support via libmtmd (#12898) |
| 1764 | 17512a94d636c4b6c1332370acb3e5af3ca70918 | 611aa914ef4231fab5d1ad04773c42e119ae2d2e | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-05-09T16:34:08+01:00 | GitHub | noreply@github.com | 2025-05-09T16:34:08+01:00 | | sycl : implementation of reordered Q4_0 MMVQ for Intel GPUs (#12858) |
| 1765 | 611aa914ef4231fab5d1ad04773c42e119ae2d2e | 0cf6725e9f9a164c39f7a87214d60342f7f946d8 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-09T15:14:56+03:00 | GitHub | noreply@github.com | 2025-05-09T15:14:56+03:00 | | metal : optimize MoE for large batches (#13388) |
| 1766 | 0cf6725e9f9a164c39f7a87214d60342f7f946d8 | 27ebfcacbaadc6104e2b18acd8f13515cbf63dce | Johannes Gäßler | johannesg@5d6.de | 2025-05-09T13:34:58+02:00 | GitHub | noreply@github.com | 2025-05-09T13:34:58+02:00 | | CUDA: FA support for Deepseek (Ampere or newer) (#13306) |
| 1767 | 27ebfcacbaadc6104e2b18acd8f13515cbf63dce | 5c86c9ed3ef1cc7307fdce05f0f0e2e45253cf90 | Diego Devesa | slarengh@gmail.com | 2025-05-09T13:02:07+02:00 | GitHub | noreply@github.com | 2025-05-09T13:02:07+02:00 | | llama : do not crash if there is no CPU backend (#13395) |
| 1768 | 5c86c9ed3ef1cc7307fdce05f0f0e2e45253cf90 | efb8b47eda78ea8ae570d4fece3953aae499289e | Johannes Gäßler | johannesg@5d6.de | 2025-05-09T12:14:04+02:00 | GitHub | noreply@github.com | 2025-05-09T12:14:04+02:00 | | CUDA: fix crash on large batch size for MoE models (#13384) |
| 1769 | efb8b47eda78ea8ae570d4fece3953aae499289e | 0527771dd80bd18479dfaaa0a98be297fc3592bf | Bartowski | 3266127+bartowski1182@users.noreply.github.com | 2025-05-09T05:53:58-04:00 | GitHub | noreply@github.com | 2025-05-09T11:53:58+02:00 | | imatrix : Add --parse-special for enabling parsing of special tokens in imatrix calculation (#13389) |
| 1770 | 0527771dd80bd18479dfaaa0a98be297fc3592bf | 2189fd3b6327a1d17893694125da8edcf74a6468 | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-05-09T17:25:50+08:00 | GitHub | noreply@github.com | 2025-05-09T10:25:50+01:00 | | llama-run: add support for downloading models from ModelScope (#13370) |
| 1771 | 2189fd3b6327a1d17893694125da8edcf74a6468 | 3f96aeff394e9b72bbd2fa665c3e023a70ed8648 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-09T11:18:02+02:00 | GitHub | noreply@github.com | 2025-05-09T11:18:02+02:00 | | mtmd : fix batch_view for m-rope (#13397) |
| 1772 | 3f96aeff394e9b72bbd2fa665c3e023a70ed8648 | b486ba05bf973aae3652b3fd593e0b257d3a41d4 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-09T11:17:51+02:00 | GitHub | noreply@github.com | 2025-05-09T11:17:51+02:00 | | llama : one-off chat template fix for Mistral-Small-2503 (#13398) |
| 1773 | b486ba05bf973aae3652b3fd593e0b257d3a41d4 | 02115dcd9a6d0c03b527d5bee8c1493ab819ddbd | Radoslav Gerganov | rgerganov@gmail.com | 2025-05-09T10:31:07+03:00 | GitHub | noreply@github.com | 2025-05-09T10:31:07+03:00 | | rpc : add rpc_msg_set_tensor_hash_req (#13353) |
| 1774 | 02115dcd9a6d0c03b527d5bee8c1493ab819ddbd | d9c4accaff30926e8a5fd8a2429e549158a845e0 | Jeff Bolz | jbolz@nvidia.com | 2025-05-09T02:23:41-05:00 | GitHub | noreply@github.com | 2025-05-09T09:23:41+02:00 | | vulkan: Allow up to 4096 elements for mul_mat_id row_ids (#13326) |
| 1775 | d9c4accaff30926e8a5fd8a2429e549158a845e0 | 15e03282bb432631193464100c2237a3b6bcfe4c | Xuan-Son Nguyen | son@huggingface.co | 2025-05-09T09:06:37+02:00 | GitHub | noreply@github.com | 2025-05-09T09:06:37+02:00 | | server : (webui) rename has_multimodal --> modalities (#13393) |
| 1776 | 15e03282bb432631193464100c2237a3b6bcfe4c | f05a6d71a0f3dbf0730b56a1abbad41c0f42e63d | Diego Devesa | slarengh@gmail.com | 2025-05-08T23:45:22+02:00 | GitHub | noreply@github.com | 2025-05-08T23:45:22+02:00 | | ci : limit write permission to only the release step + fixes (#13392) |
| 1777 | f05a6d71a0f3dbf0730b56a1abbad41c0f42e63d | ee01d71e585f4a17ae83ec55d77acfdcf0bfa798 | Matt Clayton | 156335168+mattjcly@users.noreply.github.com | 2025-05-08T14:25:39-04:00 | GitHub | noreply@github.com | 2025-05-08T20:25:39+02:00 | | mtmd : Expose helper_decode_image_chunk (#13366) |
| 1778 | ee01d71e585f4a17ae83ec55d77acfdcf0bfa798 | 8c83449cb780c201839653812681c3a4cf17feed | Xuan-Son Nguyen | son@huggingface.co | 2025-05-08T18:51:45+02:00 | GitHub | noreply@github.com | 2025-05-08T18:51:45+02:00 | | server : (webui) fix a very small misalignment (#13387) |
| 1779 | 8c83449cb780c201839653812681c3a4cf17feed | 1a844be132dbe865358e1de81136920c1d35ac73 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-08T15:37:29+02:00 | GitHub | noreply@github.com | 2025-05-08T15:37:29+02:00 | | server : (webui) revamp the input area, plus many small UI improvements (#13365) |
| 1780 | 1a844be132dbe865358e1de81136920c1d35ac73 | 0ccc1213549e39ef4c1affb1bf5f49651ef4ce48 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-08T15:34:29+02:00 | GitHub | noreply@github.com | 2025-05-08T15:34:29+02:00 | | convert : support rope_scaling type and rope_type (#13349) |
| 1781 | 0ccc1213549e39ef4c1affb1bf5f49651ef4ce48 | 6562e5a4d6c58326dcd79002ea396d4141f1b18e | welix | taichitary@gmail.com | 2025-05-08T22:03:53+09:00 | GitHub | noreply@github.com | 2025-05-08T15:03:53+02:00 | | mtmd : fix the calculation of n_tokens for smolvlm (#13381) |
| 1782 | 6562e5a4d6c58326dcd79002ea396d4141f1b18e | 51fb96b1ff2e1cc98b2492a012b7d93531a6a9a8 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-08T14:28:33+03:00 | GitHub | noreply@github.com | 2025-05-08T14:28:33+03:00 | | context : allow cache-less context for embeddings (#13108) |
| 1783 | 51fb96b1ff2e1cc98b2492a012b7d93531a6a9a8 | 70a6991edf1b60e7afa8962f830320583f3babb0 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-08T14:26:50+03:00 | GitHub | noreply@github.com | 2025-05-08T14:26:50+03:00 | | context : remove logits_all flag (#13284) |
| 1784 | 70a6991edf1b60e7afa8962f830320583f3babb0 | f0610212069929f92d428ae3b5596bfbe69c020b | Diego Devesa | slarengh@gmail.com | 2025-05-08T13:15:28+02:00 | GitHub | noreply@github.com | 2025-05-08T13:15:28+02:00 | | ci : move release workflow to a separate file (#13362) |
| 1785 | f0610212069929f92d428ae3b5596bfbe69c020b | 8733e0cf6eefc7c7752297cc22d0836706f4222c | Diego Devesa | slarengh@gmail.com | 2025-05-08T13:15:15+02:00 | GitHub | noreply@github.com | 2025-05-08T13:15:15+02:00 | | llama : print size and type of overridden tensors (#13364) |
| 1786 | 8733e0cf6eefc7c7752297cc22d0836706f4222c | 814f795e063c257f33b921eab4073484238a151a | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-05-08T10:08:01+01:00 | GitHub | noreply@github.com | 2025-05-08T10:08:01+01:00 | | sycl: addressing non-contiguous src1 mul_mats (nc and batched) (#13343) |
| 1787 | 814f795e063c257f33b921eab4073484238a151a | d879433824ac3c16b3b5f00075895d0ee9688e34 | Diego Devesa | slarengh@gmail.com | 2025-05-07T16:36:33+02:00 | GitHub | noreply@github.com | 2025-05-07T16:36:33+02:00 | | docker : disable arm64 and intel images (#13356) |
| 1788 | d879433824ac3c16b3b5f00075895d0ee9688e34 | 13b0a04597a4581cad4d9027a848f450c623801d | Georgi Gerganov | ggerganov@gmail.com | 2025-05-07T16:39:36+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-07T17:28:36+03:00 | | sync : ggml |
| 1789 | 13b0a04597a4581cad4d9027a848f450c623801d | bba9d945c14b93c5264b7956e00736601ca6f89a | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-05-05T13:09:35+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-07T17:28:36+03:00 | | whisper: remove MSVC warnings pragmas (whisper/3090) |
| 1790 | bba9d945c14b93c5264b7956e00736601ca6f89a | bc4e1128f78be0fbb4e2fa630adb6a04b969ac68 | Jared Tweed | jaredtwe@gmail.com | 2025-05-02T02:41:35-07:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-07T17:28:36+03:00 | | cmake : removed stdc++fs (whisper/3097) |
| 1791 | bc4e1128f78be0fbb4e2fa630adb6a04b969ac68 | 39e73ae0d69f882d7e29cecc6dd8f5052fca6731 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-07T12:49:27+02:00 | GitHub | noreply@github.com | 2025-05-07T12:49:27+02:00 | | llama : deci : support ffn-free with attention (#13296) |
| 1792 | 39e73ae0d69f882d7e29cecc6dd8f5052fca6731 | 1f73301b63668d61b5f3109489050a27dc3f65be | Ycros | 18012+ycros@users.noreply.github.com | 2025-05-07T18:23:28+10:00 | GitHub | noreply@github.com | 2025-05-07T11:23:28+03:00 | | common : Add a warning when we can't match samplers from a string or char. (#13330) |
| 1793 | 1f73301b63668d61b5f3109489050a27dc3f65be | 4773d7a02ffdb05ba9e673ff21ce95351836e33a | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-05-07T15:48:23+08:00 | GitHub | noreply@github.com | 2025-05-07T09:48:23+02:00 | | cuda : remove nrows_x in mul_mat_q_process_tile (#13325) |
| 1794 | 4773d7a02ffdb05ba9e673ff21ce95351836e33a | 6c7fd67b647a76846d1691cd181011dff4549d02 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-07T10:28:02+03:00 | GitHub | noreply@github.com | 2025-05-07T10:28:02+03:00 | | examples : remove infill (#13283) |
| 1795 | 6c7fd67b647a76846d1691cd181011dff4549d02 | 141a908a59bbc68ceae3bf090b850e33322a2ca9 | piDack | 104877312+piDack@users.noreply.github.com | 2025-05-07T15:23:11+08:00 | GitHub | noreply@github.com | 2025-05-07T09:23:11+02:00 | | llama : support tie embedding for chatglm models (#13328) |
| 1796 | 141a908a59bbc68ceae3bf090b850e33322a2ca9 | 32916a49072f01c43e20df374af5f8a1f70d6963 | Johannes Gäßler | johannesg@5d6.de | 2025-05-06T23:35:51+02:00 | GitHub | noreply@github.com | 2025-05-06T23:35:51+02:00 | | CUDA: mix virt/real CUDA archs for GGML_NATIVE=OFF (#13135) |
| 1797 | 32916a49072f01c43e20df374af5f8a1f70d6963 | ffc727203af1061fdeb49efef30f76171722e403 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-06T22:40:24+02:00 | GitHub | noreply@github.com | 2025-05-06T22:40:24+02:00 | | clip : refactor graph builder (#13321) |
| 1798 | ffc727203af1061fdeb49efef30f76171722e403 | 91a86a6f354aa73a7aab7bc3d283be410fdc93a5 | DocShotgun | 126566557+DocShotgun@users.noreply.github.com | 2025-05-06T13:36:24-07:00 | GitHub | noreply@github.com | 2025-05-06T22:36:24+02:00 | | sampling : make top_n_sigma no-op at <=0 or a single candidate (#13345) |
| 1799 | 91a86a6f354aa73a7aab7bc3d283be410fdc93a5 | f4ed10b69cc38c54070a47f841827de5e8984cdf | oobabooga | oobabooga4@gmail.com | 2025-05-06T15:24:15-03:00 | GitHub | noreply@github.com | 2025-05-06T20:24:15+02:00 | | sampling : don't consider -infinity values in top_n_sigma (#13344) |
| 1800 | f4ed10b69cc38c54070a47f841827de5e8984cdf | 1e333d5bba18e99bc328bb87ac1ee6a4e6260e0e | Diego Devesa | slarengh@gmail.com | 2025-05-06T20:15:31+02:00 | GitHub | noreply@github.com | 2025-05-06T20:15:31+02:00 | | cmake : remove arm64 msvc presets (#13342) |
| 1801 | 1e333d5bba18e99bc328bb87ac1ee6a4e6260e0e | 2f54e348ad2999c4e31b8777592247622b20420f | Akarshan Biswas | akarshan@menlo.ai | 2025-05-06T20:27:06+05:30 | GitHub | noreply@github.com | 2025-05-06T20:27:06+05:30 | | SYCL: Disable reorder optimize by default and stop setting tensor extras when optimize is disabled (#13254) |
| 1802 | 2f54e348ad2999c4e31b8777592247622b20420f | 2356fb1d53c86d838756211010bbabfafda7cb94 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-06T14:25:40+02:00 | GitHub | noreply@github.com | 2025-05-06T14:25:40+02:00 | | llama : fix build_ffn without gate (#13336) |
| 1803 | 2356fb1d53c86d838756211010bbabfafda7cb94 | 764b85627b46f43d7ea801867cd1b6abef484574 | Johannes Gäßler | johannesg@5d6.de | 2025-05-06T13:58:51+02:00 | GitHub | noreply@github.com | 2025-05-06T13:58:51+02:00 | | CUDA: fix bad asserts for partial offload (#13337) |
| 1804 | 764b85627b46f43d7ea801867cd1b6abef484574 | 15a28ec8c705b188ebe178170966d1dcc36fe151 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-06T11:12:06+02:00 | GitHub | noreply@github.com | 2025-05-06T11:12:06+02:00 | | convert : qwen2/3moe : set yarn metadata if present (#13331) |
| 1805 | 15a28ec8c705b188ebe178170966d1dcc36fe151 | a7366faa5bb2fff97b9fb43340d853709f52d8c9 | Johannes Gäßler | johannesg@5d6.de | 2025-05-06T08:36:46+02:00 | GitHub | noreply@github.com | 2025-05-06T08:36:46+02:00 | | CUDA: fix --split-mode row for MMQ (#13323) |
| 1806 | a7366faa5bb2fff97b9fb43340d853709f52d8c9 | 907036502070ba608bdb2aaebf802092d4cfba07 | compilade | git@compilade.net | 2025-05-05T22:27:31-04:00 | GitHub | noreply@github.com | 2025-05-05T22:27:31-04:00 | | gguf-py : avoid requiring pyside6 for other scripts (#13036) |
| 1807 | 907036502070ba608bdb2aaebf802092d4cfba07 | 233461f8121455f957a47e6a22a77b3bc88277b0 | Johannes Gäßler | johannesg@5d6.de | 2025-05-05T22:32:13+02:00 | GitHub | noreply@github.com | 2025-05-05T22:32:13+02:00 | | CUDA: fix logic for clearing padding with -ngl 0 (#13320) |
| 1808 | 233461f8121455f957a47e6a22a77b3bc88277b0 | b34c859146630dff136943abc9852ca173a7c9d6 | oobabooga | oobabooga4@gmail.com | 2025-05-05T17:12:19-03:00 | GitHub | noreply@github.com | 2025-05-05T22:12:19+02:00 | | sampling : Integrate Top-nσ into main sampling chain (and add it to the server) (#13264) |
| 1809 | b34c859146630dff136943abc9852ca173a7c9d6 | 9b61acf06041dcbaff6afa5f28940e93297f8520 | igardev | 49397134+igardev@users.noreply.github.com | 2025-05-05T17:03:31+03:00 | GitHub | noreply@github.com | 2025-05-05T16:03:31+02:00 | | server : Webui - change setText command from parent window to also send the message. (#13309) |
| 1810 | 9b61acf06041dcbaff6afa5f28940e93297f8520 | 5215b91e9377ce23e4ccc92ec3156bf5c7f892a3 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-05T16:02:55+02:00 | GitHub | noreply@github.com | 2025-05-05T16:02:55+02:00 | | mtmd : rename llava directory to mtmd (#13311) |
| 1811 | 5215b91e9377ce23e4ccc92ec3156bf5c7f892a3 | ae803bfc3d0fc2d0d3e1cce22ee103a30939e104 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-05T12:54:44+02:00 | GitHub | noreply@github.com | 2025-05-05T12:54:44+02:00 | | clip : fix confused naming ffn_up and ffn_down (#13290) |
| 1812 | ae803bfc3d0fc2d0d3e1cce22ee103a30939e104 | 66645a5285d8c4c5f9a3b3f360d042baac2d820a | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-05T12:34:26+02:00 | GitHub | noreply@github.com | 2025-05-05T12:34:26+02:00 | | convert : bailingmoe : set yarn metadata if present (#13312) |
| 1813 | 66645a5285d8c4c5f9a3b3f360d042baac2d820a | 27aa2595321c4d9cc4086a8e67bdea204b8309b0 | Akarshan Biswas | akarshan@menlo.ai | 2025-05-05T13:39:10+05:30 | GitHub | noreply@github.com | 2025-05-05T13:39:10+05:30 | | SYCL: Disable mul_mat kernels for noncontiguous tensor b (#13308) |
| 1814 | 27aa2595321c4d9cc4086a8e67bdea204b8309b0 | 9fdfcdaeddd1ef57c6d041b89cd8fb7048a0f028 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-04T23:43:42+02:00 | GitHub | noreply@github.com | 2025-05-04T23:43:42+02:00 | | mtmd : add C public API (#13184) |
| 1815 | 9fdfcdaeddd1ef57c6d041b89cd8fb7048a0f028 | 6eb7d25c700e29d9272064db408855178d01741b | Diego Devesa | slarengh@gmail.com | 2025-05-04T21:25:43+02:00 | GitHub | noreply@github.com | 2025-05-04T21:25:43+02:00 | | rpc : use backend registry, support dl backends (#13304) |
| 1816 | 6eb7d25c700e29d9272064db408855178d01741b | 86bd60d3fe4a08b8c9d920e6defbc2412d803569 | Aaron Teo | aaron.teo1@ibm.com | 2025-05-05T01:49:12+08:00 | GitHub | noreply@github.com | 2025-05-04T19:49:12+02:00 | | ggml : activate s390x simd for Q3_K (#13301) |
| 1817 | 86bd60d3fe4a08b8c9d920e6defbc2412d803569 | 9f2da5871f4bbd205b8a3b952cdc76283218d595 | Diego Devesa | slarengh@gmail.com | 2025-05-04T17:05:20+02:00 | GitHub | noreply@github.com | 2025-05-04T17:05:20+02:00 | | llava/mtmd : fixes to fully support dl backends (#13303) |
| 1818 | 9f2da5871f4bbd205b8a3b952cdc76283218d595 | 93c4e23905987949b714b21ae918ff6bfb55fe36 | Diego Devesa | slarengh@gmail.com | 2025-05-04T14:20:49+02:00 | GitHub | noreply@github.com | 2025-05-04T14:20:49+02:00 | | llama : build windows releases with dl backends (#13220) |
| 1819 | 93c4e23905987949b714b21ae918ff6bfb55fe36 | 8afbd968182909cf93fb15959fc867b6dd3adb53 | Johannes Gäßler | johannesg@5d6.de | 2025-05-04T14:16:39+02:00 | GitHub | noreply@github.com | 2025-05-04T14:16:39+02:00 | | CUDA: fix race condition in MMQ stream-k fixup (#13299) |
| 1820 | 8afbd968182909cf93fb15959fc867b6dd3adb53 | 8ae5ebcf859b05a2ea3bbd930133a2fe4a89ed3c | Johannes Gäßler | johannesg@5d6.de | 2025-05-04T13:58:38+02:00 | GitHub | noreply@github.com | 2025-05-04T13:58:38+02:00 | | CUDA: fix race condition in MMQ ids_dst (#13294) |
| 1821 | 8ae5ebcf859b05a2ea3bbd930133a2fe4a89ed3c | 3e959f09764a2bb0e64af594eab83f7fb3e08eb2 | Jeff Bolz | jbolz@nvidia.com | 2025-05-04T00:17:16-05:00 | GitHub | noreply@github.com | 2025-05-04T07:17:16+02:00 | | vulkan: Additional type support for unary, binary, and copy (#13266) |
| 1822 | 3e959f09764a2bb0e64af594eab83f7fb3e08eb2 | 36667c8edcded08063ed51c7d57e9e086bbfc903 | Johannes Gäßler | johannesg@5d6.de | 2025-05-04T00:50:37+02:00 | GitHub | noreply@github.com | 2025-05-04T00:50:37+02:00 | | imatrix: fix oob writes if src1 is not contiguous (#13286) |
| 1823 | 36667c8edcded08063ed51c7d57e9e086bbfc903 | 3bf785f3efa89ed28294fbf73054558a2b034bfb | Xuan-Son Nguyen | son@huggingface.co | 2025-05-03T20:07:54+02:00 | GitHub | noreply@github.com | 2025-05-03T20:07:54+02:00 | | clip : revert the change of BOI/EOI token for GLM-edge (⚠️ breaking change) (#13259) |
| 1824 | 3bf785f3efa89ed28294fbf73054558a2b034bfb | 1d36b3670b285e69e58b9d687c770a2a0a192194 | ymcki | 84055651+ymcki@users.noreply.github.com | 2025-05-03T23:39:51+08:00 | GitHub | noreply@github.com | 2025-05-03T17:39:51+02:00 | | llama : Llama-3_1-Nemotron-Ultra-253B-v1 support (#12843) |
| 1825 | 1d36b3670b285e69e58b9d687c770a2a0a192194 | b34443923cad751483cc53af2e680d595daadce7 | Diego Devesa | slarengh@gmail.com | 2025-05-02T20:27:13+02:00 | GitHub | noreply@github.com | 2025-05-02T20:27:13+02:00 | | llama : move end-user examples to tools directory (#13249) |
| 1826 | b34443923cad751483cc53af2e680d595daadce7 | a75cb30dc9e63488c3614e2d5a9fe2306eaf47cd | Georgi Gerganov | ggerganov@gmail.com | 2025-05-02T20:54:30+03:00 | GitHub | noreply@github.com | 2025-05-02T20:54:30+03:00 | | sync : ggml (#13268) |
| 1827 | a75cb30dc9e63488c3614e2d5a9fe2306eaf47cd | 3f3769ba76061a511f02f2a48da2ad2d93fce511 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-02T20:54:13+03:00 | GitHub | noreply@github.com | 2025-05-02T20:54:13+03:00 | | context : fix reorder logic (#13267) |
| 1828 | 3f3769ba76061a511f02f2a48da2ad2d93fce511 | 2f567611c0234bbca0a4009762acb47b56866095 | shalinib-ibm | Shalini.Salomi.Bodapati@ibm.com | 2025-05-02T22:23:12+05:30 | GitHub | noreply@github.com | 2025-05-02T19:53:12+03:00 | | ggml : Enable MMA for BF16 in llamafile_sgemm (#13148) |
| 1829 | 2f567611c0234bbca0a4009762acb47b56866095 | 7d2123484e6ba8fcd90fff8c01661b9dbbefdc9d | Jared Van Bortel | jared@nomic.ai | 2025-05-02T11:42:30-04:00 | GitHub | noreply@github.com | 2025-05-02T11:42:30-04:00 | | llama-model : support Qwen2 embedding models and pooling_mode_lasttoken (#13245) |
| 1830 | 7d2123484e6ba8fcd90fff8c01661b9dbbefdc9d | 074e42ab31de4c99aa7d9d2d239660f64b2380d6 | Jared Van Bortel | jared@nomic.ai | 2025-05-02T11:41:54-04:00 | GitHub | noreply@github.com | 2025-05-02T11:41:54-04:00 | | convert : use correct context length for nomic-embed-text-v2 (#13216) |
| 1831 | 074e42ab31de4c99aa7d9d2d239660f64b2380d6 | c642bc014c105728ce45015813351dc5a37f60a2 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-02T17:17:15+02:00 | GitHub | noreply@github.com | 2025-05-02T17:17:15+02:00 | | convert : converting mmproj for Qwen2/2.5VL from convert_hf_to_gguf (#13209) |
| 1832 | c642bc014c105728ce45015813351dc5a37f60a2 | cb06a3c363f50cd35113984fe8fb164aea419077 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-02T17:48:36+03:00 | GitHub | noreply@github.com | 2025-05-02T17:48:36+03:00 | | kv-cache : separate recurrent vs non-recurrent impl (#12799) |
| 1833 | cb06a3c363f50cd35113984fe8fb164aea419077 | 626083faf73faa54440f934bb1741bf443be91b1 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-02T12:44:24+02:00 | GitHub | noreply@github.com | 2025-05-02T12:44:24+02:00 | | llama : orion rope type is neox (#13261) |
| 1834 | 626083faf73faa54440f934bb1741bf443be91b1 | 2af6880178b4bc2c0eced726bab68b4bf333042b | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-05-02T12:40:56+02:00 | GitHub | noreply@github.com | 2025-05-02T12:40:56+02:00 | | llama : plamo rope type is neox (#13260) |
| 1835 | 2af6880178b4bc2c0eced726bab68b4bf333042b | e84773ab604fe0d935d03741df003716398dc57b | piDack | 104877312+piDack@users.noreply.github.com | 2025-05-02T17:06:09+08:00 | GitHub | noreply@github.com | 2025-05-02T11:06:09+02:00 | | llama-chat : reset glmedge chat template (#13253) |
| 1836 | e84773ab604fe0d935d03741df003716398dc57b | fab647e8842c5f80da7e8f2c625dab6a0e19e5d4 | Shakil Ahmed | 44522075+ahmedshakill@users.noreply.github.com | 2025-05-02T14:20:27+06:00 | GitHub | noreply@github.com | 2025-05-02T10:20:27+02:00 | | mtmd-cli : fix out_of_range when input image path is empty (#13244) |
| 1837 | fab647e8842c5f80da7e8f2c625dab6a0e19e5d4 | dcf886007de4b8e5200f461a13233315f897fb9d | Georgi Gerganov | ggerganov@gmail.com | 2025-05-02T09:48:31+03:00 | GitHub | noreply@github.com | 2025-05-02T09:48:31+03:00 | | server : add cache reuse card link to help (#13230) |
| 1838 | dcf886007de4b8e5200f461a13233315f897fb9d | d24d5928086471063fa9d9fd45aca710fd1336ae | Xuan-Son Nguyen | son@huggingface.co | 2025-05-02T08:45:10+02:00 | GitHub | noreply@github.com | 2025-05-02T08:45:10+02:00 | | convert : explicitly disable trust_remote_code for AutoConfig (#13246) |
| 1839 | d24d5928086471063fa9d9fd45aca710fd1336ae | 8efbdadc616fa717c369059b9b388160958d886c | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-05-01T19:06:39-03:00 | GitHub | noreply@github.com | 2025-05-01T19:06:39-03:00 | | ci: fix cross-compile sync issues (#12804) |
| 1840 | 8efbdadc616fa717c369059b9b388160958d886c | f057808ffadce213f54c1884ab5096e60140f358 | Justin Santa Barbara | justinsb@google.com | 2025-05-01T17:32:11-04:00 | GitHub | noreply@github.com | 2025-05-01T23:32:11+02:00 | | rpc : avoid uninitialized memory in serialize_tensor (#13210) |
| 1841 | f057808ffadce213f54c1884ab5096e60140f358 | d7a14c42a1883a34a6553cbfe30da1e1b84dfd6a | Jesse Gross | jesse@kernel.org | 2025-05-01T13:46:10-07:00 | GitHub | noreply@github.com | 2025-05-01T22:46:10+02:00 | | ggml: Don't assert fail when tensor data changes (#13222) |
| 1842 | d7a14c42a1883a34a6553cbfe30da1e1b84dfd6a | b6e4ff69b8abd509647b531bd5b4e86950204f66 | Diego Devesa | slarengh@gmail.com | 2025-05-01T21:48:08+02:00 | GitHub | noreply@github.com | 2025-05-01T21:48:08+02:00 | | build : fix build info on windows (#13239) |
| 1843 | b6e4ff69b8abd509647b531bd5b4e86950204f66 | e0f572c8466e70d35cbd70ee536ad8fc83b2acac | Loïc Carrère | loic.carrere@gmail.com | 2025-05-01T21:32:21+02:00 | GitHub | noreply@github.com | 2025-05-01T21:32:21+02:00 | | clip : (minicpmv) Re-enable upscaling of images smaller than the CLIP image size (#13237) |
| 1844 | e0f572c8466e70d35cbd70ee536ad8fc83b2acac | 79f26e9e125b21760aeb016f34bfd42a93f48351 | matteo | matteo.serva@gmail.com | 2025-05-01T21:16:38+02:00 | GitHub | noreply@github.com | 2025-05-01T21:16:38+02:00 | | llama-chat : update GLM4 chat template (#13238) |
| 1845 | 79f26e9e125b21760aeb016f34bfd42a93f48351 | fc727bcdd5a311c7c69a76dbf87f4784e828c7b4 | Jeff Bolz | jbolz@nvidia.com | 2025-05-01T13:49:39-05:00 | GitHub | noreply@github.com | 2025-05-01T20:49:39+02:00 | | vulkan: Add bfloat16 support (#12554) |
| 1846 | fc727bcdd5a311c7c69a76dbf87f4784e828c7b4 | b0ecbd434be024c06bf547491be444ed92e1123e | Jeff Bolz | jbolz@nvidia.com | 2025-05-01T13:19:31-05:00 | GitHub | noreply@github.com | 2025-05-01T20:19:31+02:00 | | vulkan: Handle src1 batch dimension in non-contiguous mat-vec-mul shader (#13191) |
| 1847 | b0ecbd434be024c06bf547491be444ed92e1123e | b1dd4d08e8fe40c9484422c0c5ec9bbd13b067a9 | Johannes Gäßler | johannesg@5d6.de | 2025-05-01T20:18:56+02:00 | GitHub | noreply@github.com | 2025-05-01T20:18:56+02:00 | | test: non-cont. b in test-backend-ops -o MUL_MAT (#13187) |
| 1848 | b1dd4d08e8fe40c9484422c0c5ec9bbd13b067a9 | 99881f77d82efda80b21057d84f1cc2df2f1e0f6 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T17:07:13+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T20:15:34+03:00 | | sync : ggml |
| 1849 | 99881f77d82efda80b21057d84f1cc2df2f1e0f6 | b5769d92b4510c77691ad9e3f8b643c2ba202e53 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-05-01T10:05:24+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T20:15:34+03:00 | | whisper : add check that target name exists (whisper/3103) |
| 1850 | b5769d92b4510c77691ad9e3f8b643c2ba202e53 | 8936784f7a1ec4f91637d04b77fdc90ec36ebac9 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-04-29T15:47:55+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T20:15:34+03:00 | | ggml : suppress Windows compiler warnings (whisper/3075) |
| 1851 | 8936784f7a1ec4f91637d04b77fdc90ec36ebac9 | 13c9a3319b65469e883c49dd1c478abedc410157 | Xuan-Son Nguyen | son@huggingface.co | 2025-05-01T17:05:42+02:00 | GitHub | noreply@github.com | 2025-05-01T17:05:42+02:00 | | mtmd : add **vision** support for Mistral Small 3.1 (#13231) |
| 1852 | 13c9a3319b65469e883c49dd1c478abedc410157 | a70183eb0079fdf6ebeaed12966338a039461b5d | Xuan-Son Nguyen | son@huggingface.co | 2025-05-01T10:23:25+02:00 | GitHub | noreply@github.com | 2025-05-01T10:23:25+02:00 | | arg : remove CURLINFO_EFFECTIVE_METHOD (#13228) |
| 1853 | a70183eb0079fdf6ebeaed12966338a039461b5d | 8d33d740c308fe0049515668aac0e561269afe9e | Jared Van Bortel | jared@nomic.ai | 2025-05-01T03:09:41-04:00 | GitHub | noreply@github.com | 2025-05-01T10:09:41+03:00 | | llama-model : fix the reported size class for nomic-embed-text-v2-moe (#13223) |
| 1854 | 8d33d740c308fe0049515668aac0e561269afe9e | 4254bb49518e0f920f14c9aadd96eefdfd38b429 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T09:59:02+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T10:00:39+03:00 | | sync : ggml |
| 1855 | 4254bb49518e0f920f14c9aadd96eefdfd38b429 | 9998540149a490200894acfe595e9a1546bab723 | Diego Devesa | slarengh@gmail.com | 2025-04-30T15:20:40+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T09:58:44+03:00 | | ggml : fix ggml_gallocr_ptr type (ggml/1205) |
| 1856 | 9998540149a490200894acfe595e9a1546bab723 | e1e8e0991ffd9e99a445c6812bb519d5bac9f4b5 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T18:59:06+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-05-01T09:58:44+03:00 | | cuda : fix unused variable compile warning (whisper/0) |
| 1857 | e1e8e0991ffd9e99a445c6812bb519d5bac9f4b5 | 6f67cf1f480926391ad75ff746e0a021214bf70c | Johannes Gäßler | johannesg@5d6.de | 2025-04-30T23:12:59+02:00 | GitHub | noreply@github.com | 2025-04-30T23:12:59+02:00 | | CUDA: batched+noncont MMQ, refactor bs>1 MoE code (#13199) |
| 1858 | 6f67cf1f480926391ad75ff746e0a021214bf70c | 16a457facd996915652f6274384c87602b27d21a | Xuan-Son Nguyen | son@huggingface.co | 2025-04-30T22:29:15+02:00 | GitHub | noreply@github.com | 2025-04-30T21:29:15+01:00 | | arg : -hf do not fail if url mismatch (#13219) |
| 1859 | 16a457facd996915652f6274384c87602b27d21a | 3e168bede4d27b35656ab8026015b87659ecbec2 | ddh0 | dylanhalladay02@icloud.com | 2025-04-30T15:28:43-05:00 | GitHub | noreply@github.com | 2025-04-30T21:28:43+01:00 | | fix typo: `n_ctx_pre_seq` -> `n_ctx_per_seq` (#13221) |
| 1860 | 3e168bede4d27b35656ab8026015b87659ecbec2 | ceda28ef8e310a8dee60bf275077a3eedae8e36c | Xuan-Son Nguyen | son@huggingface.co | 2025-04-30T16:56:24+02:00 | GitHub | noreply@github.com | 2025-04-30T16:56:24+02:00 | | convert : improve model arch handling (#13122) |
| 1861 | ceda28ef8e310a8dee60bf275077a3eedae8e36c | 3b127c738535d95e06abd0d43da147bc13516ad0 | Tatsuya Tanaka | tanakasan2525@gmail.com | 2025-04-30T22:25:20+09:00 | GitHub | noreply@github.com | 2025-04-30T15:25:20+02:00 | | llava : remove duplicate include (#13207) |
| 1862 | 3b127c738535d95e06abd0d43da147bc13516ad0 | e5007a5edf2692ef7151a81a61ce2716b83374e5 | Olivier Chafik | ochafik@users.noreply.github.com | 2025-04-30T13:52:35+01:00 | GitHub | noreply@github.com | 2025-04-30T14:52:35+02:00 | | common : add -jf / --json-schema-file flag (#12011) |
| 1863 | e5007a5edf2692ef7151a81a61ce2716b83374e5 | 416313773b53585fddcafbcb914cbbfbaeb94b1f | Jeff Bolz | jbolz@nvidia.com | 2025-04-30T07:38:37-05:00 | GitHub | noreply@github.com | 2025-04-30T14:38:37+02:00 | | vulkan: use uint array index to avoid glslang bug (#13193) |
| 1864 | 416313773b53585fddcafbcb914cbbfbaeb94b1f | 07c2e2f76cce9a61c110b6995fbb90ccea2c3aaa | shalinib-ibm | Shalini.Salomi.Bodapati@ibm.com | 2025-04-30T16:47:08+05:30 | GitHub | noreply@github.com | 2025-04-30T13:17:08+02:00 | | ggml : fix ppc64le build (#13176) |
| 1865 | 07c2e2f76cce9a61c110b6995fbb90ccea2c3aaa | 44cd8d91ff2c9e4a0f2e3151f8d6f04c928e2571 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-30T13:06:15+02:00 | GitHub | noreply@github.com | 2025-04-30T13:06:15+02:00 | | convert : correct typo image_mean --> image_std (#13208) |
| 1866 | 44cd8d91ff2c9e4a0f2e3151f8d6f04c928e2571 | 5933e6fdc9c051eea6c83b5a7608de12f9f15670 | Aaron Teo | 57927438+taronaeo@users.noreply.github.com | 2025-04-30T17:47:35+08:00 | GitHub | noreply@github.com | 2025-04-30T10:47:35+01:00 | | feat(ggml-cpu): enable z17 compile (#13182) |
| 1867 | 5933e6fdc9c051eea6c83b5a7608de12f9f15670 | da84c04d8fa43ff92b172feb8130c74d062f956a | Xuan-Son Nguyen | son@huggingface.co | 2025-04-30T10:46:32+02:00 | GitHub | noreply@github.com | 2025-04-30T10:46:32+02:00 | | arg : allow using -hf offline (#13202) |
| 1868 | da84c04d8fa43ff92b172feb8130c74d062f956a | a0f7016d170ca4bfe24d9a9f26c024d034af69f2 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-30T10:44:07+02:00 | GitHub | noreply@github.com | 2025-04-30T10:44:07+02:00 | | docker : do not build tests (#13204) |
| 1869 | a0f7016d170ca4bfe24d9a9f26c024d034af69f2 | 19e899ce21a7c9ffcf8bb2b22269a75f6e078f8f | xiaofei | hbuxiaofei@gmail.com | 2025-04-30T14:29:22+08:00 | GitHub | noreply@github.com | 2025-04-30T09:29:22+03:00 | | rpc : fix cache directory initialization (#13188) |
| 1870 | 19e899ce21a7c9ffcf8bb2b22269a75f6e078f8f | e2e1ddb93a01ce282e304431b37e60b3cddb6114 | Johannes Gäßler | johannesg@5d6.de | 2025-04-29T23:32:04+02:00 | GitHub | noreply@github.com | 2025-04-29T23:32:04+02:00 | | scripts: n_depth for compare-llama-bench [no ci] (#13201) |
| 1871 | e2e1ddb93a01ce282e304431b37e60b3cddb6114 | d9d398f84f96d16c308d4976f5e90222ecc2a492 | matteo | matteo.serva@gmail.com | 2025-04-29T20:33:10+02:00 | GitHub | noreply@github.com | 2025-04-29T20:33:10+02:00 | | server : Prefilling assistant message in openai compatible API (#13174) |
| 1872 | d9d398f84f96d16c308d4976f5e90222ecc2a492 | 5a6398011704c31178d7b774be67856ba57647c8 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-29T20:22:57+03:00 | GitHub | noreply@github.com | 2025-04-29T20:22:57+03:00 | | sampling : when top-k <= 0 -> noop (#13173) |
| 1873 | 5a6398011704c31178d7b774be67856ba57647c8 | cdf76586b23c67abd3ca064ee2c084c57ae240bd | Alberto Cabrera Pérez | alberto.cabrera@codeplay.com | 2025-04-29T16:24:36+01:00 | GitHub | noreply@github.com | 2025-04-29T17:24:36+02:00 | | llama-bench: fixed size of fields to correctly map to values (#13183) |
| 1874 | cdf76586b23c67abd3ca064ee2c084c57ae240bd | 7d3af70b089bb349b5d17eb01839224c99ec1952 | Johannes Gäßler | johannesg@5d6.de | 2025-04-29T16:00:27+02:00 | GitHub | noreply@github.com | 2025-04-29T16:00:27+02:00 | | CUDA: fix non-cont. inputs for batched mat mul (#13155) |
| 1875 | 7d3af70b089bb349b5d17eb01839224c99ec1952 | 00e3e5a194e88e604e7c91391b9e90332888fd72 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-04-29T13:25:53+02:00 | GitHub | noreply@github.com | 2025-04-29T13:25:53+02:00 | | llama : llm_type order by size (#13177) |
| 1876 | 00e3e5a194e88e604e7c91391b9e90332888fd72 | e98b3692be4cd8fbbd9a56fbacc2f2bf0bf26a68 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-29T11:47:04+02:00 | GitHub | noreply@github.com | 2025-04-29T11:47:04+02:00 | | mtmd : add qwen2vl and qwen2.5vl (#13141) |
| 1877 | e98b3692be4cd8fbbd9a56fbacc2f2bf0bf26a68 | b6ce7430b7eb51f032152316880204e0a9c0470e | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-04-29T11:00:31+02:00 | GitHub | noreply@github.com | 2025-04-29T11:00:31+02:00 | | llama : set qwen3 model type sizes (#13175) |
| 1878 | b6ce7430b7eb51f032152316880204e0a9c0470e | 5f5e39e1ba5dbea814e41f2a15e035d749a520bc | Xuan-Son Nguyen | son@huggingface.co | 2025-04-29T08:45:49+02:00 | GitHub | noreply@github.com | 2025-04-29T09:45:49+03:00 | | llama-graph : fix text position for mrope (#13159) |
| 1879 | 5f5e39e1ba5dbea814e41f2a15e035d749a520bc | eaea3253244dc4bbe07f6cd81325847ccc6cf93e | AT | manyoso@users.noreply.github.com | 2025-04-28T15:52:15-04:00 | GitHub | noreply@github.com | 2025-04-28T22:52:15+03:00 | | model : Nomic Embed Text V2 with Mixture-of-Experts (MoE) architecture (#12466) |
| 1880 | eaea3253244dc4bbe07f6cd81325847ccc6cf93e | 43ddab6eeeaab5a04fe5a364af0bafb0e4d35065 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T21:23:19+02:00 | GitHub | noreply@github.com | 2025-04-28T21:23:19+02:00 | | clip : fix model size display (#13153) |
| 1881 | 43ddab6eeeaab5a04fe5a364af0bafb0e4d35065 | 1831f538f720d1d99fba146f24f0a8e970838cc4 | Ville Vesilehto | ville@vesilehto.fi | 2025-04-28T21:00:20+03:00 | GitHub | noreply@github.com | 2025-04-28T21:00:20+03:00 | | fix(rpc): Improve input validation and error handling (#13069) |
| 1882 | 1831f538f720d1d99fba146f24f0a8e970838cc4 | 4e87962e34a4b257ec374c4baf6b1568554b81a9 | Vishal Agarwal | vishalagarwal.jss@gmail.com | 2025-04-28T20:20:39+05:30 | GitHub | noreply@github.com | 2025-04-28T16:50:39+02:00 | | llama-bench: add `-d` depth arg (#13096) |
| 1883 | 4e87962e34a4b257ec374c4baf6b1568554b81a9 | fb0471d1753824e75474c24f82fbdd54c94dceda | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T16:12:56+02:00 | GitHub | noreply@github.com | 2025-04-28T16:12:56+02:00 | | mtmd : fix glm-edge redundant token count (#13139) |
| 1884 | fb0471d1753824e75474c24f82fbdd54c94dceda | d2b2031e5f11b826dcc718138642f147a2009665 | pockers21 | 134406831+pockers21@users.noreply.github.com | 2025-04-28T06:45:40-07:00 | GitHub | noreply@github.com | 2025-04-28T16:45:40+03:00 | | context : do not clear output buffer on reserve (#13152) |
| 1885 | d2b2031e5f11b826dcc718138642f147a2009665 | 5fa9e63be82225fb3249c76f39ddda3e5bdec0a3 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T14:20:56+02:00 | GitHub | noreply@github.com | 2025-04-28T14:20:56+02:00 | | llama : (mrope) allow using normal 1D position for text token (#13138) |
| 1886 | 5fa9e63be82225fb3249c76f39ddda3e5bdec0a3 | a4c340f974f9b7ac0c1aae897aabaa54549a97e5 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T12:18:59+02:00 | GitHub | noreply@github.com | 2025-04-28T12:18:59+02:00 | | clip : refactor set input for cgraph + fix qwen2.5vl input (#13136) |
| 1887 | a4c340f974f9b7ac0c1aae897aabaa54549a97e5 | d0a417f3c7a5a22ef05b3b76d91dbe1d3362bf0c | Akarshan Biswas | akarshan@menlo.ai | 2025-04-28T15:03:25+05:30 | GitHub | noreply@github.com | 2025-04-28T11:33:25+02:00 | | SYCL: Add all missing unary kernels (#13074) |
| 1888 | d0a417f3c7a5a22ef05b3b76d91dbe1d3362bf0c | 43f2b07193cbcccd266734320ea9b948f5a01926 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-28T12:10:18+03:00 | GitHub | noreply@github.com | 2025-04-28T12:10:18+03:00 | | readme : update hot topics (#13150) |
| 1889 | 43f2b07193cbcccd266734320ea9b948f5a01926 | e5d6c2554e7597665e26991a93fa2f3d16c79ad5 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-28T11:57:19+03:00 | GitHub | noreply@github.com | 2025-04-28T11:57:19+03:00 | | common : fix noreturn compile warning (#13151) |
| 1890 | e5d6c2554e7597665e26991a93fa2f3d16c79ad5 | f0dd6a1926cdb2f4183a937deee40db26ef8f1da | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T10:11:58+02:00 | GitHub | noreply@github.com | 2025-04-28T10:11:58+02:00 | | llama-chat : fix typo GML --> GLM (#13143) |
| 1891 | f0dd6a1926cdb2f4183a937deee40db26ef8f1da | 69699be48a6b94570773532850667f1591dc5bbe | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-28T15:33:28+08:00 | GitHub | noreply@github.com | 2025-04-28T09:33:28+02:00 | | musa: fix typo in cc control (#13144) |
| 1892 | 69699be48a6b94570773532850667f1591dc5bbe | 85f36e5e7173eef7c671c778db44c034e1d0ab19 | Johannes Gäßler | johannesg@5d6.de | 2025-04-28T09:29:26+02:00 | GitHub | noreply@github.com | 2025-04-28T09:29:26+02:00 | | CUDA: fix q_nope_absorbed prec for DS 2 Lite f16 (#13137) |
| 1893 | 85f36e5e7173eef7c671c778db44c034e1d0ab19 | c0a97b762e5ec767dc414f0dc4979befd4c09a52 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-28T07:16:59+02:00 | GitHub | noreply@github.com | 2025-04-28T08:16:59+03:00 | | arg : fix unused variable (#13142) |
| 1894 | c0a97b762e5ec767dc414f0dc4979befd4c09a52 | ced44be34290fab450f8344efa047d8a08e723b4 | 4onen | 11580688+4onen@users.noreply.github.com | 2025-04-27T14:48:26-07:00 | GitHub | noreply@github.com | 2025-04-27T23:48:26+02:00 | | llama-bench : Add `--override-tensors` arg (#12922) |
| 1895 | ced44be34290fab450f8344efa047d8a08e723b4 | e291450b7602d7a36239e4ceeece37625f838373 | matteo | matteo.serva@gmail.com | 2025-04-27T21:57:32+02:00 | GitHub | noreply@github.com | 2025-04-27T21:57:32+02:00 | | llama-chat : fix wrong template in GLM4-0414 (#13140) |
| 1896 | e291450b7602d7a36239e4ceeece37625f838373 | 59e991c23cc44cb5fb657e7b9358cac21fb79828 | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-27T19:22:49+08:00 | GitHub | noreply@github.com | 2025-04-27T13:22:49+02:00 | | musa: fix build warning (#13129) |
| 1897 | 59e991c23cc44cb5fb657e7b9358cac21fb79828 | ca2bb89eac2097ab4620448737e58af8452e444b | LostRuins Concedo | 39025047+LostRuins@users.noreply.github.com | 2025-04-27T18:43:37+08:00 | GitHub | noreply@github.com | 2025-04-27T12:43:37+02:00 | | Fixes Qwen2.5VL segfault during inference with https://github.com/ggml-org/llama.cpp/pull/12402 as has_qwen2vl_merger migration was incomplete (#13133) |
| 1898 | ca2bb89eac2097ab4620448737e58af8452e444b | 2d451c80590b9ac250322769ac13d3b4870dbcf7 | HimariO | dsfhe49854@gmail.com | 2025-04-27T16:10:34+08:00 | GitHub | noreply@github.com | 2025-04-27T10:10:34+02:00 | | clip : Add Qwen2.5VL support (#12402) |
| 1899 | 2d451c80590b9ac250322769ac13d3b4870dbcf7 | 4753791e70acd4d4e02f2098f14a03df26c992bd | Xuan-Son Nguyen | son@huggingface.co | 2025-04-26T22:58:12+02:00 | GitHub | noreply@github.com | 2025-04-26T22:58:12+02:00 | | common : add common_remote_get_content (#13123) |
| 1900 | 4753791e70acd4d4e02f2098f14a03df26c992bd | 77d5e9a76a7b4a8a7c5bf9cf6ebef91860123cba | Xuan-Son Nguyen | son@huggingface.co | 2025-04-26T22:39:47+02:00 | GitHub | noreply@github.com | 2025-04-26T22:39:47+02:00 | | clip : improve projector naming (#13118) |
| 1901 | 77d5e9a76a7b4a8a7c5bf9cf6ebef91860123cba | d5fe4e81bd447124836ecfb47d794f8768665b9f | SXX | sxx1136965276@gmail.com | 2025-04-26T22:05:31+08:00 | GitHub | noreply@github.com | 2025-04-26T16:05:31+02:00 | | ggml: move fp16/bf16 conversion optimizations to CPU backend + export conversion APIs (#13107) |
| 1902 | d5fe4e81bd447124836ecfb47d794f8768665b9f | 295354ea6848a77bdee204ee1c971d9b92ffcca9 | frob | rick+github@frob.com.au | 2025-04-26T10:10:20+02:00 | GitHub | noreply@github.com | 2025-04-26T10:10:20+02:00 | | grammar : handle maxItems == 0 in JSON schema (#13117) |
| 1903 | 295354ea6848a77bdee204ee1c971d9b92ffcca9 | 558a764713468f26f5a163d25a22100c9a04a48f | Diego Devesa | slarengh@gmail.com | 2025-04-25T19:40:11+02:00 | GitHub | noreply@github.com | 2025-04-25T19:40:11+02:00 | | llama : fix K-shift with quantized K and BLAS backend (#13113) |
| 1904 | 558a764713468f26f5a163d25a22100c9a04a48f | edb18b6e8f5ea6509ad43057f8bb98fc557dbc4e | City | 125218114+city96@users.noreply.github.com | 2025-04-25T14:38:34+02:00 | GitHub | noreply@github.com | 2025-04-25T14:38:34+02:00 | | Force FP32 compute in GLM4 FFN Down (#13101) |
| 1905 | edb18b6e8f5ea6509ad43057f8bb98fc557dbc4e | 514c45608f93f66106a712dee1abe062099ce790 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-25T14:31:42+02:00 | GitHub | noreply@github.com | 2025-04-25T14:31:42+02:00 | | clip : fix pixtral on some GPU backends (#13097) |
| 1906 | 514c45608f93f66106a712dee1abe062099ce790 | 553a5c3a9fdf771be2101bc3529937963f817457 | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-04-25T17:37:51+08:00 | GitHub | noreply@github.com | 2025-04-25T17:37:51+08:00 | | change the reorder tensor from init to execute OP (#13003) |
| 1907 | 553a5c3a9fdf771be2101bc3529937963f817457 | 13be08daf992c89d5169518229b3740041c0f419 | Radoslav Gerganov | rgerganov@gmail.com | 2025-04-25T10:08:08+03:00 | GitHub | noreply@github.com | 2025-04-25T10:08:08+03:00 | | rpc : do not wait for response when sending RPC_CMD_SET_TENSOR (#12943) |
| 1908 | 13be08daf992c89d5169518229b3740041c0f419 | 226251ed56b85190e18a1cca963c45b888f4953c | Xuan-Son Nguyen | son@huggingface.co | 2025-04-24T22:17:04+02:00 | GitHub | noreply@github.com | 2025-04-24T22:17:04+02:00 | | clip : remove boi/eoi embeddings for GLM-edge model (#13081) |
| 1909 | 226251ed56b85190e18a1cca963c45b888f4953c | 87616f0680947800ecba3e9f6bc6e101943bf8e6 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T22:29:22+03:00 | GitHub | noreply@github.com | 2025-04-24T22:29:22+03:00 | | embeddings : fix batch sizes (#13076) |
| 1910 | 87616f0680947800ecba3e9f6bc6e101943bf8e6 | 63b4911494afe04778c61b9c19019341d71c99fc | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T17:22:27+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T17:32:47+03:00 | | ggml : fix trailing whitespaces (#0) |
| 1911 | 63b4911494afe04778c61b9c19019341d71c99fc | c6e8cc28c15166dba15629dba6a7366d4d5955ca | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T16:47:43+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T17:32:47+03:00 | | sync : ggml |
| 1912 | c6e8cc28c15166dba15629dba6a7366d4d5955ca | b10d8bfdb1dac40cce34e8860ca5ec7d950c3a44 | Acly | aclysia@gmail.com | 2025-04-17T14:16:45+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T17:32:47+03:00 | | ggml : Depthwise 2D convolution (ggml/1152) |
| 1913 | b10d8bfdb1dac40cce34e8860ca5ec7d950c3a44 | 13b4548877326fdabee3e831b8cfd65d9844383c | Johannes Gäßler | johannesg@5d6.de | 2025-04-24T15:57:10+02:00 | GitHub | noreply@github.com | 2025-04-24T15:57:10+02:00 | | CUDA: use switch statements in constexpr functions (#13095) |
| 1914 | 13b4548877326fdabee3e831b8cfd65d9844383c | 572b3141d343d7f947bf53b57513016e90db5680 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T16:00:10+03:00 | GitHub | noreply@github.com | 2025-04-24T16:00:10+03:00 | | cmake : do not include ./src as public for libllama (#13062) |
| 1915 | 572b3141d343d7f947bf53b57513016e90db5680 | 7c727fbe39150fbe8381f4fa43fed08719ebebe6 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T15:44:05+03:00 | GitHub | noreply@github.com | 2025-04-24T15:44:05+03:00 | | clang-tidy : disable warning about missing math parenthesis (#13091) |
| 1916 | 7c727fbe39150fbe8381f4fa43fed08719ebebe6 | 80982e815e67bae2442237f4e11466f44c9a2988 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-24T14:04:14+02:00 | GitHub | noreply@github.com | 2025-04-24T14:04:14+02:00 | | arg : add --no-mmproj-offload (#13093) |
| 1917 | 80982e815e67bae2442237f4e11466f44c9a2988 | 7604a7d6b80e78eef8f275fc700d0f64820d672f | Xuan-Son Nguyen | son@huggingface.co | 2025-04-24T12:14:13+02:00 | GitHub | noreply@github.com | 2025-04-24T12:14:13+02:00 | | arg : clean up handling --mmproj with -hf (#13082) |
| 1918 | 7604a7d6b80e78eef8f275fc700d0f64820d672f | b3b6d862cfdf190e1b9ad961639a25f5ebc0c7e3 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-24T10:38:30+03:00 | GitHub | noreply@github.com | 2025-04-24T10:38:30+03:00 | | metal : fix floating-point range of attention scores in FA kernels (#13090) |
| 1919 | b3b6d862cfdf190e1b9ad961639a25f5ebc0c7e3 | 56304069599f4dd9749d94b9bca2c2c65bb27c02 | Eve | 139727413+netrunnereve@users.noreply.github.com | 2025-04-24T07:18:33Z | GitHub | noreply@github.com | 2025-04-24T09:18:33+02:00 | | vulkan: matmul gcn tuning (#13016) |
| 1920 | 56304069599f4dd9749d94b9bca2c2c65bb27c02 | ecda2ec4b347031a9b8a89ee2efc664ce63f599c | pl752 | pl752@mail.ru | 2025-04-24T02:32:35+05:00 | GitHub | noreply@github.com | 2025-04-23T23:32:35+02:00 | | llama-mtmd-cli: Sigint rework in mtmd vision example (#13080) |
| 1921 | ecda2ec4b347031a9b8a89ee2efc664ce63f599c | eb1776b15a32d832f1266deeeab75b9d255c5849 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-23T20:21:59+02:00 | GitHub | noreply@github.com | 2025-04-23T20:21:59+02:00 | | mtmd : Support Pixtral 12B (#13065) |
| 1922 | eb1776b15a32d832f1266deeeab75b9d255c5849 | 2cca6c01e46d2fc1124d15730273ed2acdad1016 | piDack | 104877312+piDack@users.noreply.github.com | 2025-04-23T22:59:14+08:00 | GitHub | noreply@github.com | 2025-04-23T16:59:14+02:00 | | convert : Append mult-eos,half-rope,bos to GLM4-0414 and Z (#13021) |
| 1923 | 2cca6c01e46d2fc1124d15730273ed2acdad1016 | 658987cfc9d752dca7758987390d5fb1a7a0a54a | Radoslav Gerganov | rgerganov@gmail.com | 2025-04-23T10:32:49+03:00 | GitHub | noreply@github.com | 2025-04-23T10:32:49+03:00 | | rpc : add command line option for number of threads for the CPU backend (#13060) |
| 1924 | 658987cfc9d752dca7758987390d5fb1a7a0a54a | dc39a5e7a84815a90fa0c515ed8927870cf858c9 | Johannes Gäßler | johannesg@5d6.de | 2025-04-22T21:27:40+02:00 | GitHub | noreply@github.com | 2025-04-22T21:27:40+02:00 | | CUDA: noncont MMVQ + batched bs1 MUL_MAT_ID (#13014) |
| 1925 | dc39a5e7a84815a90fa0c515ed8927870cf858c9 | ab47dec3d37aa1927c2ec590e166b76141374ed3 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-22T16:24:54+02:00 | GitHub | noreply@github.com | 2025-04-22T16:24:54+02:00 | | mtmd : support SmolVLM (version 1 and 2) (#13050) |
| 1926 | ab47dec3d37aa1927c2ec590e166b76141374ed3 | 7b53389c24a507564be39e1ea82746a39749059b | Georgi Gerganov | ggerganov@gmail.com | 2025-04-22T16:16:10+03:00 | GitHub | noreply@github.com | 2025-04-22T16:16:10+03:00 | | security : add note about RPC and server functionality (#13061) |
| 1927 | 7b53389c24a507564be39e1ea82746a39749059b | 243453533e029334181dda50d911d5fc5a2b2486 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-22T16:15:51+03:00 | GitHub | noreply@github.com | 2025-04-22T16:15:51+03:00 | | metal : add memory pool for temp allocs (#12850) |
| 1928 | 243453533e029334181dda50d911d5fc5a2b2486 | 1d735c0b4fa0551c51c2f4ac888dd9a01f447985 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-22T10:37:00+02:00 | GitHub | noreply@github.com | 2025-04-22T10:37:00+02:00 | | llava : update documentations (#13055) |
| 1929 | 1d735c0b4fa0551c51c2f4ac888dd9a01f447985 | 5368ddda7a262d195b54687a31009dcc1f8b1602 | Diego Devesa | slarengh@gmail.com | 2025-04-21T18:13:51+02:00 | GitHub | noreply@github.com | 2025-04-21T18:13:51+02:00 | | ggml : add SSE 4.2 and x64 base variant for CPUs without AVX (#12871) |
| 1930 | 5368ddda7a262d195b54687a31009dcc1f8b1602 | 84a9bf2fc2875205f0806fbbfbb66dc67204094c | Akarshan Biswas | akarshan@menlo.ai | 2025-04-21T19:13:30+05:30 | GitHub | noreply@github.com | 2025-04-21T19:13:30+05:30 | | SYCL: Add non-contiguous support in ROPE (#12993) |
| 1931 | 84a9bf2fc2875205f0806fbbfbb66dc67204094c | 2016f07bd106c73699ecbaace80f55db5ed95dac | Xuan-Son Nguyen | son@huggingface.co | 2025-04-21T15:32:58+02:00 | GitHub | noreply@github.com | 2025-04-21T15:32:58+02:00 | | mtmd : merge llava, gemma3 and minicpmv CLI into single `llama-mtmd-cli` (#13012) |
| 1932 | 2016f07bd106c73699ecbaace80f55db5ed95dac | 6602304814e679cc8c162bb760a034aceb4f8965 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-20T23:29:36+02:00 | GitHub | noreply@github.com | 2025-04-20T23:29:36+02:00 | | convert : experimental support for `--mmproj` flag (#13023) |
| 1933 | 6602304814e679cc8c162bb760a034aceb4f8965 | 66168204be9559dc841f06f0025f3da01d9a8546 | Jeffrey Morgan | jmorganca@gmail.com | 2025-04-20T03:15:41-07:00 | GitHub | noreply@github.com | 2025-04-20T12:15:41+02:00 | | llava: fix errors in clip.h on certain compilers (#13030) |
| 1934 | 66168204be9559dc841f06f0025f3da01d9a8546 | 4ba9d711ba0a5bf0be00936d3258d4a5e4164d4f | Jeff Bolz | jbolz@nvidia.com | 2025-04-20T03:50:02-05:00 | GitHub | noreply@github.com | 2025-04-20T10:50:02+02:00 | | vulkan: support noncontiguous rms_norm (#13031) |
| 1935 | 4ba9d711ba0a5bf0be00936d3258d4a5e4164d4f | 00137157fca3d17b90380762b4d7cc158d385bd3 | Jeffrey Morgan | jmorganca@gmail.com | 2025-04-19T22:28:40-07:00 | GitHub | noreply@github.com | 2025-04-20T08:28:40+03:00 | | metal: add neg operator (#13029) |
| 1936 | 00137157fca3d17b90380762b4d7cc158d385bd3 | fb28f4f80efb5a2990a6a240433b02e259b0dbb1 | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-04-19T13:05:03-03:00 | GitHub | noreply@github.com | 2025-04-19T18:05:03+02:00 | | Disable CI cross-compile builds (#13022) |
| 1937 | fb28f4f80efb5a2990a6a240433b02e259b0dbb1 | 37b9f0d29d7301dd0ec5dfae8f357ccee96325d7 | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-04-19T16:26:38+02:00 | GitHub | noreply@github.com | 2025-04-19T16:26:38+02:00 | | gguf-py : fix upload python package workflow (#13020) |
| 1938 | 37b9f0d29d7301dd0ec5dfae8f357ccee96325d7 | 6408210082cc0a61b992b487be7e2ff2efbb9e36 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-19T09:15:45+02:00 | GitHub | noreply@github.com | 2025-04-19T09:15:45+02:00 | | clip : refactor, add `image_manipulation` and `llava_uhd` classes (#13011) |
| 1939 | 6408210082cc0a61b992b487be7e2ff2efbb9e36 | aff9d107b066c9d464f7ab1324280acfdcebf569 | Daniel Tang | danielzgtg.opensource@gmail.com | 2025-04-18T16:02:55-04:00 | GitHub | noreply@github.com | 2025-04-18T22:02:55+02:00 | | main : Fix Ctrl+D/newline handling (#12951) |
| 1940 | aff9d107b066c9d464f7ab1324280acfdcebf569 | 35370ba94593c3abc9550328377d461fc4f822b7 | Chris Thompson | christopherthompson81@gmail.com | 2025-04-18T12:30:41-06:00 | GitHub | noreply@github.com | 2025-04-18T20:30:41+02:00 | | gguf-py : GGUF Editor GUI - Python + Qt6 (#12930) |
| 1941 | 35370ba94593c3abc9550328377d461fc4f822b7 | 8d6600576318dfc6b091fca744b0fd36a5e5255f | Xuan-Son Nguyen | son@huggingface.co | 2025-04-18T19:58:12+02:00 | GitHub | noreply@github.com | 2025-04-18T19:58:12+02:00 | | server : use std::move whenever possible (#12936) |
| 1942 | 8d6600576318dfc6b091fca744b0fd36a5e5255f | b9154ecff93ff54dc554411eb844a2a654be49f2 | Akarshan Biswas | akarshan@menlo.ai | 2025-04-18T19:27:56+05:30 | GitHub | noreply@github.com | 2025-04-18T15:57:56+02:00 | | SYCL: Refactor and enable FP16 in binary broadcast OPs (#12975) |
| 1943 | b9154ecff93ff54dc554411eb844a2a654be49f2 | 2db9ba1464f3de0aceb2b5289963e69fc369cb66 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-18T10:04:51+02:00 | GitHub | noreply@github.com | 2025-04-18T10:04:51+02:00 | | mtmd : add methods to access `mtmd_image_tokens` (#12906) |
| 1944 | 2db9ba1464f3de0aceb2b5289963e69fc369cb66 | 2f74c354c0f752ed9aabf7d3a350e6edebd7e744 | Radoslav Gerganov | rgerganov@gmail.com | 2025-04-18T10:13:42+03:00 | GitHub | noreply@github.com | 2025-04-18T10:13:42+03:00 | | rpc : add RPC_CMD_HELLO (#12955) |
| 1945 | 2f74c354c0f752ed9aabf7d3a350e6edebd7e744 | 207c22ec2d6d793fc70830138617d1e016c5151c | Georgi Gerganov | ggerganov@gmail.com | 2025-04-17T18:16:36+03:00 | GitHub | noreply@github.com | 2025-04-17T18:16:36+03:00 | | graph : make FA compatible with MLA + add initial Metal kernels (#12953) |
| 1946 | 207c22ec2d6d793fc70830138617d1e016c5151c | 7a395f67a7a02bb361d944b816d6e933889e28e1 | Alan Gray | agray3@users.noreply.github.com | 2025-04-17T14:19:42+01:00 | GitHub | noreply@github.com | 2025-04-17T15:19:42+02:00 | | ggml: Re-enable CUDA graphs in presence of CONT and DUP nodes (#12970) |
| 1947 | 7a395f67a7a02bb361d944b816d6e933889e28e1 | 971f245b3b5f3f55991bb779cb541b00f82eea1d | hipudding | huafengchun@gmail.com | 2025-04-17T20:34:16+08:00 | GitHub | noreply@github.com | 2025-04-17T20:34:16+08:00 | | CANN: Add support for async operator submission (#12864) |
| 1948 | 971f245b3b5f3f55991bb779cb541b00f82eea1d | 12b17501e6015ffe568ac54fdf08e6580833bf1b | Mikko Juola | mikjuo@gmail.com | 2025-04-17T01:37:05-07:00 | GitHub | noreply@github.com | 2025-04-17T11:37:05+03:00 | | llama : recognize IBM Granite 3.3 FIM tokens (#12988) |
| 1949 | 12b17501e6015ffe568ac54fdf08e6580833bf1b | 015022bb53387baa8b23817ac03743705c7d472b | kimminsu | 80271594+kimminsu38oo@users.noreply.github.com | 2025-04-17T06:25:57+09:00 | GitHub | noreply@github.com | 2025-04-16T14:25:57-07:00 | | opencl: fix incorrect local_size index in profiling log (#12868) |
| 1950 | 015022bb53387baa8b23817ac03743705c7d472b | b43d89e311c5e7fbf62e5ec3c0401eb536677267 | Jeff Bolz | jbolz@nvidia.com | 2025-04-16T13:37:25-05:00 | GitHub | noreply@github.com | 2025-04-16T20:37:25+02:00 | | vulkan: enable coopmat2 FA gqa and split_k optimizations more often (#12931) |
| 1951 | b43d89e311c5e7fbf62e5ec3c0401eb536677267 | 80f19b41869728eeb6a26569957b92a773a2b2c6 | Chenguang Li | 757486878@qq.com | 2025-04-16T16:21:05+08:00 | GitHub | noreply@github.com | 2025-04-16T16:21:05+08:00 | | CANN: Add 310P operator support check (#12962) |
| 1952 | 80f19b41869728eeb6a26569957b92a773a2b2c6 | f8f820cc4dc37032d5375972ba904ce53043445d | lhez | quic_lih@quicinc.com | 2025-04-15T12:26:00-07:00 | GitHub | noreply@github.com | 2025-04-15T12:26:00-07:00 | | opencl: split `ggml-opencl.cl` into multiple files and cleanup (#12886) |
| 1953 | f8f820cc4dc37032d5375972ba904ce53043445d | 54a72720432810c89c2693f34908b60b88da1e46 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-15T14:45:05+03:00 | GitHub | noreply@github.com | 2025-04-15T14:45:05+03:00 | | metal : add FA-vec kernels for head size 96 (#12952) |
| 1954 | 54a72720432810c89c2693f34908b60b88da1e46 | 84778e97703740d8ac5fb64e14d83b80eafa0f3c | hipudding | huafengchun@gmail.com | 2025-04-15T19:08:55+08:00 | GitHub | noreply@github.com | 2025-04-15T12:08:55+01:00 | | CANN: Add x86 build ci (#12950) |
| 1955 | 84778e97703740d8ac5fb64e14d83b80eafa0f3c | 510676475f885ec064ff147af9f20ee7a9b12a50 | David Huang | 1969802+hjc4869@users.noreply.github.com | 2025-04-15T17:20:38+08:00 | GitHub | noreply@github.com | 2025-04-15T11:20:38+02:00 | | CUDA/HIP: Share the same unified memory allocation logic. (#12934) |
| 1956 | 510676475f885ec064ff147af9f20ee7a9b12a50 | daa422881a0ec7944771bcc8ff8de34d11f5bd3b | Akarshan Biswas | akarshan@menlo.ai | 2025-04-15T14:07:42+05:30 | GitHub | noreply@github.com | 2025-04-15T10:37:42+02:00 | | SYCL: Add ROPE vision kernel (#12887) |
| 1957 | daa422881a0ec7944771bcc8ff8de34d11f5bd3b | eccc7a1602f0752507de4aaad1008b9618a282c8 | Juk Armstrong | 69222624+jukofyork@users.noreply.github.com | 2025-04-15T07:49:57+01:00 | GitHub | noreply@github.com | 2025-04-15T09:49:57+03:00 | | llama : DeepSeek V2/V3 MLA implementation (#12801) |
| 1958 | eccc7a1602f0752507de4aaad1008b9618a282c8 | 0019279bb56f028e4eee19b59b19750736c719f7 | Srihari-mcw | 96763064+Srihari-mcw@users.noreply.github.com | 2025-04-15T11:52:36+05:30 | GitHub | noreply@github.com | 2025-04-15T09:22:36+03:00 | | ggml : Add AVX512 implementation of GEMM - Q4_Kx8 (#12829) |
| 1959 | 0019279bb56f028e4eee19b59b19750736c719f7 | b0c75ac9f93322b45e3766e149417334b4fd1ed9 | Chenguang Li | 757486878@qq.com | 2025-04-15T10:09:35+08:00 | GitHub | noreply@github.com | 2025-04-15T10:09:35+08:00 | | CANN: Opt ROPE optimization (#12865) |
| 1960 | b0c75ac9f93322b45e3766e149417334b4fd1ed9 | d6d2c2ab8c8865784ba9fef37f2b2de3f2134d33 | Xinpeng Dou | 15529241576@163.com | 2025-04-15T10:04:24+08:00 | GitHub | noreply@github.com | 2025-04-15T10:04:24+08:00 | | CANN: Optimize CANN buffer pool memory management (#12875) |
| 1961 | d6d2c2ab8c8865784ba9fef37f2b2de3f2134d33 | 75afa0ae31f0a51aaadcc5ff146eb7a32a7f9088 | Russyyds | 161207317+Russyyds@users.noreply.github.com | 2025-04-15T01:18:20+08:00 | GitHub | noreply@github.com | 2025-04-14T19:18:20+02:00 | | Add performance print for gemma3 in example (#12929) |
| 1962 | 75afa0ae31f0a51aaadcc5ff146eb7a32a7f9088 | c772d549264c1be058411312a54049e0dc86a037 | Akarshan Biswas | akarshan@menlo.ai | 2025-04-14T17:53:53+05:30 | GitHub | noreply@github.com | 2025-04-14T14:23:53+02:00 | | SYCL: Fix im2col (#12910) |
| 1963 | c772d549264c1be058411312a54049e0dc86a037 | 81c7e64fc239288e91a58adad9145110e0353822 | Radoslav Gerganov | rgerganov@gmail.com | 2025-04-14T13:59:34+03:00 | GitHub | noreply@github.com | 2025-04-14T13:59:34+03:00 | | rpc : use ggml_context_ptr (#12938) |
| 1964 | 81c7e64fc239288e91a58adad9145110e0353822 | 526739b879c9d6a702ed681138611197fa4d3f18 | Neo Zhang Jianyu | jianyu.zhang@intel.com | 2025-04-14T18:19:07+08:00 | GitHub | noreply@github.com | 2025-04-14T18:19:07+08:00 | | dsiable curl lib check, this action is missed by commit bd3f59f81289b920bcc597a208c14f55e39ed37e (#12761) (#12937) |
| 1965 | 526739b879c9d6a702ed681138611197fa4d3f18 | a25355e2643b1feb2fcbbc4c00312aa86370d858 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-14T08:52:10+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-14T09:26:15+03:00 | | sync : ggml |
| 1966 | a25355e2643b1feb2fcbbc4c00312aa86370d858 | e959d32b1c16c02c0bc0bcdc7bc3f7077bc32f99 | cmdr2 | secondary.cmdr2@gmail.com | 2025-04-11T12:14:19+05:30 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-14T09:26:15+03:00 | | cpu: fix cpu backend's supports-op for GET_ROWS_BACK. fixes a fatal when running test-backend-ops with only the CPU backend (ggml/1190) |
| 1967 | e959d32b1c16c02c0bc0bcdc7bc3f7077bc32f99 | 307bfa253dea07c9270e78fa53b133504e9c3c9d | SXX | sxx1136965276@gmail.com | 2025-04-14T13:47:55+08:00 | GitHub | noreply@github.com | 2025-04-14T08:47:55+03:00 | | ggml: use _mm[512/256]_dpbusd[_avx]_epi32 to directly accumulate into the result register (#12773) |
| 1968 | 307bfa253dea07c9270e78fa53b133504e9c3c9d | 71e90e8813f90097701e62f7fce137d96ddf41e2 | Alan Gray | agray3@users.noreply.github.com | 2025-04-13T22:12:21+01:00 | GitHub | noreply@github.com | 2025-04-13T23:12:21+02:00 | | ggml: disable CUDA graphs for unsupported DUP and CONT node types (#12891) |
| 1969 | 71e90e8813f90097701e62f7fce137d96ddf41e2 | bc091a4dc585af25c438c8473285a8cfec5c7695 | Ed Addario | 29247825+EAddario@users.noreply.github.com | 2025-04-13T19:29:28+01:00 | GitHub | noreply@github.com | 2025-04-13T21:29:28+03:00 | | quantize: Handle user-defined quantization levels for additional tensors (#12511) |
| 1970 | bc091a4dc585af25c438c8473285a8cfec5c7695 | a4837577aae59c0ae640f3810094724bcac6bb28 | Prajwal B Mehendarkar | prajwal.b.mehendarkar@ibm.com | 2025-04-12T21:03:39+05:30 | GitHub | noreply@github.com | 2025-04-12T17:33:39+02:00 | | common : Define cache directory on AIX (#12915) |
| 1971 | a4837577aae59c0ae640f3810094724bcac6bb28 | e59ea539b83d2c7947c99bd350549364dbba450c | Jeff Bolz | jbolz@nvidia.com | 2025-04-12T03:44:48-05:00 | GitHub | noreply@github.com | 2025-04-12T10:44:48+02:00 | | vulkan: use aligned loads for flash attention mask (#12853) |
| 1972 | e59ea539b83d2c7947c99bd350549364dbba450c | c94085df2871d2cad11f6411304d7798d7dd4cdd | Matt Clayton | 156335168+mattjcly@users.noreply.github.com | 2025-04-12T01:29:03-04:00 | GitHub | noreply@github.com | 2025-04-12T07:29:03+02:00 | | llava: Fix cpu-only clip image encoding sefault (#12907) |
| 1973 | c94085df2871d2cad11f6411304d7798d7dd4cdd | e8a62631b3b05bcf2ad0e0c686881a9f3e3f03ca | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T23:37:41+03:00 | GitHub | noreply@github.com | 2025-04-11T23:37:41+03:00 | | server : add VSCode's Github Copilot Chat support (#12896) |
| 1974 | e8a62631b3b05bcf2ad0e0c686881a9f3e3f03ca | b6930ebc421e20399663bbd3bc7b7266d5236d06 | yuri@FreeBSD | yurivict@users.noreply.github.com | 2025-04-11T13:04:14-07:00 | GitHub | noreply@github.com | 2025-04-11T22:04:14+02:00 | | rpc : Set cache directory in rpc-server.cpp on FreeBSD (#12903) |
| 1975 | b6930ebc421e20399663bbd3bc7b7266d5236d06 | 68b08f36d047bfa048ebd7d16262c53bf9ece701 | Olivier Chafik | ochafik@users.noreply.github.com | 2025-04-11T12:47:52-07:00 | GitHub | noreply@github.com | 2025-04-11T21:47:52+02:00 | | `tool-call`: fix non-tool-calling grammar crashes w/ Qwen / Hermes 2 templates (#12900) |
| 1976 | 68b08f36d047bfa048ebd7d16262c53bf9ece701 | 578754b3157d662c2fdd51eaa62b6c1f43d3172c | yuri@FreeBSD | yurivict@users.noreply.github.com | 2025-04-11T12:45:44-07:00 | GitHub | noreply@github.com | 2025-04-11T21:45:44+02:00 | | common : Define cache directory on FreeBSD (#12892) |
| 1977 | 578754b3157d662c2fdd51eaa62b6c1f43d3172c | b2034c2b55b36b2192bdefb3b295db2a911370f5 | Ewan Crawford | ewan.cr@gmail.com | 2025-04-11T15:32:14+02:00 | GitHub | noreply@github.com | 2025-04-11T15:32:14+02:00 | | sycl: Support sycl_ext_oneapi_limited_graph (#12873) |
| 1978 | b2034c2b55b36b2192bdefb3b295db2a911370f5 | 06bb53ad9b6e6d92b6ab6979927530080b1c990c | tastelikefeet | 58414341+tastelikefeet@users.noreply.github.com | 2025-04-11T20:01:56+08:00 | GitHub | noreply@github.com | 2025-04-11T14:01:56+02:00 | | contrib: support modelscope community (#12664) |
| 1979 | 06bb53ad9b6e6d92b6ab6979927530080b1c990c | 0c509239445088f21265580b86733c5b184261d9 | Yuxuan Zhang | 2448370773@qq.com | 2025-04-11T18:10:10+08:00 | GitHub | noreply@github.com | 2025-04-11T12:10:10+02:00 | | llama-model : add Glm4Model implementation for GLM-4-0414 (#12867) |
| 1980 | 0c509239445088f21265580b86733c5b184261d9 | fccf9cae83e6c6cd31a0ecb403d237638e427d0a | Xuan-Son Nguyen | son@huggingface.co | 2025-04-11T12:09:39+02:00 | GitHub | noreply@github.com | 2025-04-11T12:09:39+02:00 | | clip : use smart pointer (⚠️ breaking change) (#12869) |
| 1981 | fccf9cae83e6c6cd31a0ecb403d237638e427d0a | ec6c09d0fac1f2699c3ea1994e10482cb4b95e0f | Akarshan Biswas | akarshan@menlo.ai | 2025-04-11T13:33:50+05:30 | GitHub | noreply@github.com | 2025-04-11T16:03:50+08:00 | | SYCL: Add fp16 type support to unary op kernels (#12788) |
| 1982 | ec6c09d0fac1f2699c3ea1994e10482cb4b95e0f | 8ac9f5d765f2b1b7f821873618c0cff68ca59bc8 | Daniel Han | danielhanchen@gmail.com | 2025-04-11T00:49:09-07:00 | GitHub | noreply@github.com | 2025-04-11T09:49:09+02:00 | | convert : Llama4 RoPE fix (#12889) |
| 1983 | 8ac9f5d765f2b1b7f821873618c0cff68ca59bc8 | 12e9158f25cb47e97ca9b8083e1ec7415f262260 | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-11T15:26:17+08:00 | GitHub | noreply@github.com | 2025-04-11T09:26:17+02:00 | | ci : Replace freediskspace to free_disk_space in docker.yml (#12861) |
| 1984 | 12e9158f25cb47e97ca9b8083e1ec7415f262260 | 5b1f13cb64dbb0de7e95dae86725928d350c188c | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-04-11T09:24:34+02:00 | GitHub | noreply@github.com | 2025-04-11T09:24:34+02:00 | | xcf : add check for visionos build version (#12854) |
| 1985 | 5b1f13cb64dbb0de7e95dae86725928d350c188c | 8b91d5355a84bbb5a73fd23fb658272473240efe | Xuan-Son Nguyen | son@huggingface.co | 2025-04-11T09:23:37+02:00 | GitHub | noreply@github.com | 2025-04-11T09:23:37+02:00 | | convert : proper tensor name mapping for llama4 (#12870) |
| 1986 | 8b91d5355a84bbb5a73fd23fb658272473240efe | 0fed24c34787d2aa916ed204db88889591ad6e55 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-11T08:49:50+02:00 | GitHub | noreply@github.com | 2025-04-11T08:49:50+02:00 | | llama : correct rms norm for llama 4 (#12882) |
| 1987 | 0fed24c34787d2aa916ed204db88889591ad6e55 | 47ba87d0a4f91eb181ea60f8a7fc2539b123aeaf | Aaron Teo | 57927438+taronaeo@users.noreply.github.com | 2025-04-11T13:20:07+08:00 | GitHub | noreply@github.com | 2025-04-11T08:20:07+03:00 | | ggml: fix compilation error s390x (#12848) |
| 1988 | 47ba87d0a4f91eb181ea60f8a7fc2539b123aeaf | 1d2b613445bab6b179296aa63c8ee346fb947cc4 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:08:23+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | sync : ggml |
| 1989 | 1d2b613445bab6b179296aa63c8ee346fb947cc4 | eb420e11484ef32cc1abedb2a26ecef1bbf1e31b | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:04:25+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | tests : fix init order (#0) |
| 1990 | eb420e11484ef32cc1abedb2a26ecef1bbf1e31b | cb79c2e7fa28e874694c14598c1fd2fde82263e1 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-10T23:59:16+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | sync : ggml |
| 1991 | cb79c2e7fa28e874694c14598c1fd2fde82263e1 | fe92821ea9ae53f3088cf2699a9e102448295fa0 | cmdr2 | secondary.cmdr2@gmail.com | 2025-04-10T17:53:08+05:30 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | ggml: don't include arm_neon.h when using CUDA 12 with ARM Neon (ggml/1187) |
| 1992 | fe92821ea9ae53f3088cf2699a9e102448295fa0 | 459895c32629e36cc3ca53bf3596d3e1322cfe09 | Diego Devesa | slarengh@gmail.com | 2025-04-09T12:32:13+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | ggml : add bilinear upscale support (ggml/1185) |
| 1993 | 459895c32629e36cc3ca53bf3596d3e1322cfe09 | e4bf72d631434b38656cb97e2411826d900433e8 | Diego Devesa | slarengh@gmail.com | 2025-04-09T12:31:34+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | ggml : add more generic custom op, remove deprecated custom ops (ggml/1183) |
| 1994 | e4bf72d631434b38656cb97e2411826d900433e8 | 8b9cc7cdd8a0dcf0176c60c755322c95b5965299 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-10T23:59:01+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-11T00:17:47+03:00 | | scripts : fix sync-ggml-am.sh |
| 1995 | 8b9cc7cdd8a0dcf0176c60c755322c95b5965299 | 64eda5deb9859e87a020e56bab5d2f9ca956f1de | Xuan-Son Nguyen | son@huggingface.co | 2025-04-10T22:57:16+02:00 | GitHub | noreply@github.com | 2025-04-10T22:57:16+02:00 | | llava : introduce libmtmd (#12849) |
| 1996 | 64eda5deb9859e87a020e56bab5d2f9ca956f1de | fe5b78c89670b2f37ecb216306bed3e677b49d9f | Xuan-Son Nguyen | son@huggingface.co | 2025-04-10T17:24:44+02:00 | GitHub | noreply@github.com | 2025-04-10T17:24:44+02:00 | | convert : ability to lazy-load safetensors remotely without downloading to disk (#12820) |
| 1997 | fe5b78c89670b2f37ecb216306bed3e677b49d9f | 11d07e1e69138b46375e9267b31acd58e3813577 | Chenguang Li | 757486878@qq.com | 2025-04-10T08:51:52+08:00 | GitHub | noreply@github.com | 2025-04-10T08:51:52+08:00 | | CANN: Support more ops (#12841) |
| 1998 | 11d07e1e69138b46375e9267b31acd58e3813577 | b0091ecc1e5c0f689be856fade3803a534f35c9f | Prajwal B Mehendarkar | prajwal.b.mehendarkar@ibm.com | 2025-04-10T04:48:01+05:30 | GitHub | noreply@github.com | 2025-04-10T01:18:01+02:00 | | Fixes #12823 (#12830) |
| 1999 | b0091ecc1e5c0f689be856fade3803a534f35c9f | 31f7803bc4e7c0dcc279ee04c2ecfb76b2afdd3e | Rudi Servo | rudiservo@gmail.com | 2025-04-09T23:17:12Z | GitHub | noreply@github.com | 2025-04-10T01:17:12+02:00 | | docker : added all CPU to GPU images (#12749) |
| 2000 | 31f7803bc4e7c0dcc279ee04c2ecfb76b2afdd3e | 2391506ace6abb56186def40c7107fdfa694ed55 | Piotr Kubaj | pkubaj@anongoth.pl | 2025-04-09T23:00:34Z | GitHub | noreply@github.com | 2025-04-10T01:00:34+02:00 | | ggml-cpu-impl.h: do not redefine bool on POWER9 (#12856) |
| 2001 | 2391506ace6abb56186def40c7107fdfa694ed55 | d3bd7193ba66c15963fd1c59448f22019a8caf6e | Piotr Kubaj | pkubaj@anongoth.pl | 2025-04-09T23:00:25Z | GitHub | noreply@github.com | 2025-04-10T01:00:25+02:00 | | ggml-impl.h: fix build on POWER9 (#12855) |
| 2002 | d3bd7193ba66c15963fd1c59448f22019a8caf6e | d9a63b2f2e91cdcb0eda211b7f49fadcdea0f664 | Bo Zheng | 368586905@qq.com | 2025-04-09T17:47:36+08:00 | GitHub | noreply@github.com | 2025-04-09T11:47:36+02:00 | | llama : Support Qwen3 and Qwen3MoE (#12828) |
| 2003 | d9a63b2f2e91cdcb0eda211b7f49fadcdea0f664 | 8ed71242f464dc0a3fb3cffcfe064e55bdec72f9 | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-09T17:22:30+08:00 | GitHub | noreply@github.com | 2025-04-09T11:22:30+02:00 | | musa: enable freediskspace for docker image build (#12839) |
| 2004 | 8ed71242f464dc0a3fb3cffcfe064e55bdec72f9 | 381603a77504ad1788965f694094540c1bed9ea2 | Romain Biessy | romain.biessy@codeplay.com | 2025-04-09T11:22:04+02:00 | GitHub | noreply@github.com | 2025-04-09T11:22:04+02:00 | | sycl: update documentation to use -no-cnv (#12845) |
| 2005 | 381603a77504ad1788965f694094540c1bed9ea2 | 65a69e6e1b6d55cd5f78f8bcdfaba8a8c59a8d96 | Plamen Minev | pacominev@gmail.com | 2025-04-09T11:11:11+03:00 | GitHub | noreply@github.com | 2025-04-09T10:11:11+02:00 | | ci: detach common from the library (#12827) |
| 2006 | 65a69e6e1b6d55cd5f78f8bcdfaba8a8c59a8d96 | 47277d6d1d0d515cff34292a1a78a0d1b7252350 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-09T10:09:53+02:00 | GitHub | noreply@github.com | 2025-04-09T10:09:53+02:00 | | clip : do not print ftype (#12832) |
| 2007 | 47277d6d1d0d515cff34292a1a78a0d1b7252350 | 6e1c4cebdb697f925c523d3a969128d945161bdd | Georgi Gerganov | ggerganov@gmail.com | 2025-04-09T10:54:42+03:00 | GitHub | noreply@github.com | 2025-04-09T10:54:42+03:00 | | readme : add rpc backend (#12842) |
| 2008 | 6e1c4cebdb697f925c523d3a969128d945161bdd | 0090950f679475c5ecaac2f7bca5049cca96492b | Chenguang Li | 757486878@qq.com | 2025-04-09T14:04:14+08:00 | GitHub | noreply@github.com | 2025-04-09T14:04:14+08:00 | | CANN: Support Opt CONV_TRANSPOSE_1D and ELU (#12786) |
| 2009 | 0090950f679475c5ecaac2f7bca5049cca96492b | 7ecd780b1a1d5214b8d04c25ebfc194d310816ed | Jeff Bolz | jbolz@nvidia.com | 2025-04-09T00:25:08-05:00 | GitHub | noreply@github.com | 2025-04-09T07:25:08+02:00 | | vulkan: In coopmat2 mmq, load q4_k/q5_k scales through shared memory (#12833) |
| 2010 | 7ecd780b1a1d5214b8d04c25ebfc194d310816ed | 7538246e7ce0606694c38055cc2fc9f60535be6c | Jeff Bolz | jbolz@nvidia.com | 2025-04-09T00:12:57-05:00 | GitHub | noreply@github.com | 2025-04-09T07:12:57+02:00 | | vulkan: Use fp16 for the flash attention P*V multiplication (#12783) |
| 2011 | 7538246e7ce0606694c38055cc2fc9f60535be6c | b32efad2bc42460637c3a364c9554ea8217b3d7f | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-04-08T23:21:31+02:00 | GitHub | noreply@github.com | 2025-04-08T23:21:31+02:00 | | cuda : add f32 to bf16 copy op (#12806) |
| 2012 | b32efad2bc42460637c3a364c9554ea8217b3d7f | a19b5cef16d885c44c635da4a5c97113c1577de8 | Matt Clayton | 156335168+mattjcly@users.noreply.github.com | 2025-04-08T16:01:58-04:00 | GitHub | noreply@github.com | 2025-04-08T22:01:58+02:00 | | llava: improve clip_ctx destructor to not memleak load_image_size (#12834) |
| 2013 | a19b5cef16d885c44c635da4a5c97113c1577de8 | 78a1ba0a4f2bfed5b8b8e312592143d22e531698 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-08T19:54:51+03:00 | GitHub | noreply@github.com | 2025-04-08T19:54:51+03:00 | | llama : fix FA when KV cache is not used (i.e. embeddings) (#12825) |
| 2014 | 78a1ba0a4f2bfed5b8b8e312592143d22e531698 | 2dabf759e7c8c827d38cdd18d0792070e6a4f4e1 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-08T18:37:06+02:00 | GitHub | noreply@github.com | 2025-04-08T18:37:06+02:00 | | server : fix thread.join() on exit (#12831) |
| 2015 | 2dabf759e7c8c827d38cdd18d0792070e6a4f4e1 | 1d343b4069c74b2c7b19ae84260cd98aa2320a9a | dm4 | sunrisedm4@gmail.com | 2025-04-08T21:49:13+08:00 | GitHub | noreply@github.com | 2025-04-08T15:49:13+02:00 | | llava: add more helper functions to check projector types in clip context (#12824) |
| 2016 | 1d343b4069c74b2c7b19ae84260cd98aa2320a9a | 8ca6e1c3a4deb6bb27ee294bfd5706098d94ae88 | Prajwal B Mehendarkar | prajwal.b.mehendarkar@ibm.com | 2025-04-08T18:00:59+05:30 | GitHub | noreply@github.com | 2025-04-08T14:30:59+02:00 | | arg : Including limits file on AIX (#12822) |
| 2017 | 8ca6e1c3a4deb6bb27ee294bfd5706098d94ae88 | 656babd6c21a3b9b3622324cfcc80a2ab78da25b | characharm | 123120856+characharm@users.noreply.github.com | 2025-04-08T14:14:59+05:00 | GitHub | noreply@github.com | 2025-04-08T11:14:59+02:00 | | server : webui : Improve Chat Input with Auto-Sizing Textarea (#12785) |
| 2018 | a226bc7a9ac50551f9f113808de0f0046837f188 | 1466621e738779eefe1bb672e17dc55d63d166bb | compilade | git@compilade.net | 2025-04-08T03:03:07-04:00 | GitHub | noreply@github.com | 2025-04-08T09:03:07+02:00 | | gguf-py : support lazy tensor splitting (#12809) |
| 2019 | 1466621e738779eefe1bb672e17dc55d63d166bb | 82974011f312057b446c27267105bd7ad3810599 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-07T23:06:44+02:00 | GitHub | noreply@github.com | 2025-04-07T23:06:44+02:00 | | llama : Support llama 4 text-only (#12791) |
| 2020 | 82974011f312057b446c27267105bd7ad3810599 | 4ccea213bc629c4eef7b520f7f6c59ce9bbdaca0 | lhez | quic_lih@quicinc.com | 2025-04-07T13:22:54-07:00 | GitHub | noreply@github.com | 2025-04-07T13:22:54-07:00 | | opencl: better identify Adreno GPU (#12760) |
| 2021 | 4ccea213bc629c4eef7b520f7f6c59ce9bbdaca0 | 1a1ab7e7a4a9b6e6440c1d9965d2b9d1b7e7dafb | stduhpf | stephduh@live.fr | 2025-04-07T17:47:08+02:00 | GitHub | noreply@github.com | 2025-04-07T18:47:08+03:00 | | hellaswag: display estimated score confidence interval (#12797) |
| 2022 | 1a1ab7e7a4a9b6e6440c1d9965d2b9d1b7e7dafb | a4e46e28f99d7f64f38d8328cdb6bee5a3a1cd03 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T13:18:07+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T18:44:17+03:00 | | cuda : fix HIP and MUSA BF16 (#0) |
| 2023 | a4e46e28f99d7f64f38d8328cdb6bee5a3a1cd03 | ff067dbcb9e6ca4ed464d3db999ff8e9c503498b | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T12:32:39+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T18:44:17+03:00 | | sync : ggml |
| 2024 | ff067dbcb9e6ca4ed464d3db999ff8e9c503498b | 36ca8b362885e4b4984d72e348ba403568864a95 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T12:25:15+03:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T18:44:17+03:00 | | ggml : simplify Arm fp16 CPU logic (ggml/1177) |
| 2025 | 36ca8b362885e4b4984d72e348ba403568864a95 | 995083e4ed24933e6a289472f9de0f0b53ca5eca | Sigbjørn Skjæret | sigbjorn.skjaeret@scala.com | 2025-04-04T21:05:12+02:00 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T18:44:17+03:00 | | CUDA: don't convert BF16 weights to FP32 (ggml/1174) |
| 2026 | 995083e4ed24933e6a289472f9de0f0b53ca5eca | 518a01480eb3a7c80a4951b430db9dee55428310 | cmdr2 | secondary.cmdr2@gmail.com | 2025-04-02T17:46:16+05:30 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-07T18:44:17+03:00 | | cpu: move all the operators into a separate c++ file (except mul_mat) (ggml/1167) |
| 2027 | 518a01480eb3a7c80a4951b430db9dee55428310 | e391d3ee8ddae86be70c034de1082ad51c55e211 | zhouwg | zhouwg2000@gmail.com | 2025-04-07T23:22:57+08:00 | GitHub | noreply@github.com | 2025-04-07T17:22:57+02:00 | | sycl: remove redundant memcopy in function ggml_backend_sycl_buffer_set_tensor (#12734) |
| 2028 | e391d3ee8ddae86be70c034de1082ad51c55e211 | bd3f59f81289b920bcc597a208c14f55e39ed37e | Xuan-Son Nguyen | son@huggingface.co | 2025-04-07T14:37:28+02:00 | GitHub | noreply@github.com | 2025-04-07T15:37:28+03:00 | | ci : no curl on ggml-ci (#12796) |
| 2029 | bd3f59f81289b920bcc597a208c14f55e39ed37e | 52b3d71f128c1eeab39b918adf1f7a8f5e256619 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-07T13:35:19+02:00 | GitHub | noreply@github.com | 2025-04-07T13:35:19+02:00 | | cmake : enable curl by default (#12761) |
| 2030 | 52b3d71f128c1eeab39b918adf1f7a8f5e256619 | d0d5b2232b3826cac2c6e62d0244a474ba79860f | zhouwg | zhouwg2000@gmail.com | 2025-04-07T19:34:14+08:00 | GitHub | noreply@github.com | 2025-04-07T19:34:14+08:00 | | CANN: fix typo in ggml-cann (#12733) |
| 2031 | d0d5b2232b3826cac2c6e62d0244a474ba79860f | 916c83bfe7f8b08ada609c3b8e583cf5301e594b | hipudding | huafengchun@gmail.com | 2025-04-07T17:10:36+08:00 | GitHub | noreply@github.com | 2025-04-07T17:10:36+08:00 | | CANN: Refactor to reduce duplicate code (#12731) |
| 2032 | 916c83bfe7f8b08ada609c3b8e583cf5301e594b | 0c74b04376b0b9efc096480fe10f866afc8d7c1c | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-06T21:23:54+08:00 | GitHub | noreply@github.com | 2025-04-06T15:23:54+02:00 | | musa: fix compilation warnings in mp_22/31 (#12780) |
| 2033 | 0c74b04376b0b9efc096480fe10f866afc8d7c1c | 80b717d493a9a5bae7167ad2384c12c60bb2ef20 | Jeff Bolz | jbolz@nvidia.com | 2025-04-06T04:03:47-05:00 | GitHub | noreply@github.com | 2025-04-06T11:03:47+02:00 | | vulkan: fix NaN issue in flash attention shader (#12776) |
| 2034 | 80b717d493a9a5bae7167ad2384c12c60bb2ef20 | 6bf28f0111ff9f21b3c1b1eace20c590281e7ba6 | Jeff Bolz | jbolz@nvidia.com | 2025-04-06T03:47:13-05:00 | GitHub | noreply@github.com | 2025-04-06T10:47:13+02:00 | | vulkan: Use unclamped loads for flash attention mask (#12720) |
| 2035 | 6bf28f0111ff9f21b3c1b1eace20c590281e7ba6 | f1e3eb4249db68d97be352c0adf16eef7ae53795 | 0cc4m | picard12@live.de | 2025-04-05T18:04:03+02:00 | GitHub | noreply@github.com | 2025-04-05T18:04:03+02:00 | | Vulkan: Tune Vulkan mmq int dot shader for performance (#12767) |
| 2036 | f1e3eb4249db68d97be352c0adf16eef7ae53795 | 0364178ca2452bd241a4cdd4350d0d7a6ff8c4b4 | Sergey Fedorov | vital.had@gmail.com | 2025-04-05T23:46:00+08:00 | GitHub | noreply@github.com | 2025-04-05T17:46:00+02:00 | | common : fix includes in arg.cpp and gemma3-cli.cpp (#12766) |
| 2037 | 0364178ca2452bd241a4cdd4350d0d7a6ff8c4b4 | c6ff5d2a8da2587cc78c9ede9171dfc3f076c757 | Xuan-Son Nguyen | son@huggingface.co | 2025-04-05T17:17:40+02:00 | GitHub | noreply@github.com | 2025-04-05T17:17:40+02:00 | | clip : refactor clip_init, add tests (#12757) |
| 2038 | c6ff5d2a8da2587cc78c9ede9171dfc3f076c757 | 7a84777f42a9b3ba47db5d20b7662f8ddf92f652 | エシュナヴァリシア | 148695646+eternaphia@users.noreply.github.com | 2025-04-05T21:31:42+08:00 | GitHub | noreply@github.com | 2025-04-05T15:31:42+02:00 | | common: custom hf endpoint support (#12769) |
| 2039 | 7a84777f42a9b3ba47db5d20b7662f8ddf92f652 | 3e1d29348b5d77269f6931500dd1c1a729d429c8 | Olivier Chafik | ochafik@users.noreply.github.com | 2025-04-04T13:16:39-07:00 | GitHub | noreply@github.com | 2025-04-04T21:16:39+01:00 | | sync: minja (#12739) |
| 2040 | 3e1d29348b5d77269f6931500dd1c1a729d429c8 | 1be76e4620ddc99b3cbff0f206436117ae561047 | Georgi Gerganov | ggerganov@gmail.com | 2025-04-04T21:48:10+03:00 | GitHub | noreply@github.com | 2025-04-04T21:48:10+03:00 | | kv-cache : simplify + fix warning for recurrent models (#12756) |
| 2041 | 1be76e4620ddc99b3cbff0f206436117ae561047 | b7723942970b70412793f1926a35cfb93d5c157c | bandoti | 141645996+bandoti@users.noreply.github.com | 2025-04-04T14:05:12-03:00 | GitHub | noreply@github.com | 2025-04-04T14:05:12-03:00 | | ci: add Linux cross-compile build (#12428) |
| 2042 | b7723942970b70412793f1926a35cfb93d5c157c | 23106f94ea2bc3da929afb7330655fd5515d08dc | Nauful Shaikh | nauful@gmail.com | 2025-04-04T09:09:52-05:00 | GitHub | noreply@github.com | 2025-04-04T16:09:52+02:00 | | server : webui : Upgrade daisyui, tailwindcss. (#12735) |
| 2043 | 23106f94ea2bc3da929afb7330655fd5515d08dc | 94148ba330968bbfb8d9ecc67751bdc2218486cd | nick huang | nickhuang99@hotmail.com | 2025-04-04T22:09:12+08:00 | GitHub | noreply@github.com | 2025-04-04T16:09:12+02:00 | | gguf-split : --merge now respects --dry-run option (#12681) |
| 2044 | 94148ba330968bbfb8d9ecc67751bdc2218486cd | 9ac4d611d05b03f961e627f4f399e4377bdb4ad3 | Nicolò Scipione | nicolo.scipione@codeplay.com | 2025-04-04T16:00:46+02:00 | GitHub | noreply@github.com | 2025-04-04T16:00:46+02:00 | | sycl: allow ggml-sycl configuration and compilation using Visual Studio project/solution (#12625) |
| 2045 | 9ac4d611d05b03f961e627f4f399e4377bdb4ad3 | 348888e0dc469a476f9e9eccb394893c513daab8 | Ronny Brendel | ronnyb@nvidia.com | 2025-04-04T15:12:40+02:00 | GitHub | noreply@github.com | 2025-04-04T10:12:40-03:00 | | cmake: fix ggml-shaders-gen compiler paths containing spaces (#12747) |
| 2046 | 348888e0dc469a476f9e9eccb394893c513daab8 | 74d4f5b041ad837153b0e90fc864b8290e01d8d5 | Daniel Bevenius | daniel.bevenius@gmail.com | 2025-04-04T10:24:12+02:00 | GitHub | noreply@github.com | 2025-04-04T10:24:12+02:00 | | docs : add XCFramework section to README.md [no ci] (#12746) |
| 2047 | 74d4f5b041ad837153b0e90fc864b8290e01d8d5 | 35e592eb30832187412360912ab8f2f5b7984df1 | Jeff Bolz | jbolz@nvidia.com | 2025-04-04T00:54:35-05:00 | GitHub | noreply@github.com | 2025-04-04T07:54:35+02:00 | | vulkan: Hybrid waitForFences/getFenceStatus to reduce fence latency (#12630) |
| 2048 | 35e592eb30832187412360912ab8f2f5b7984df1 | 7d7b1bafa7980603a61cb102879fe38a33f66c08 | Jeff Bolz | jbolz@nvidia.com | 2025-04-04T00:53:20-05:00 | GitHub | noreply@github.com | 2025-04-04T07:53:20+02:00 | | vulkan: set cmake minimum and project name in vulkan-shaders (#12744) |
| 2049 | 7d7b1bafa7980603a61cb102879fe38a33f66c08 | c262beddf29f3f3be5bbbf167b56029a19876956 | lhez | quic_lih@quicinc.com | 2025-04-03T22:18:17-07:00 | GitHub | noreply@github.com | 2025-04-03T22:18:17-07:00 | | opencl: update doc for OpenCL (#12702) |
| 2050 | c262beddf29f3f3be5bbbf167b56029a19876956 | 5dd5d1ab00d074e3b7c02ca3ae12f6bf3e86336a | Gaurav Garg | 52341457+gaugarg-nv@users.noreply.github.com | 2025-04-03T21:50:29+05:30 | GitHub | noreply@github.com | 2025-04-03T18:20:29+02:00 | | CUDA: Prefer vector flash decoding kernel for Gemma models (#12738) |
| 2051 | 5dd5d1ab00d074e3b7c02ca3ae12f6bf3e86336a | 1c059995e080e37aeaacaae35850fc4c044b9dbe | yumeyao | yumeyao@gmail.com | 2025-04-03T23:32:54+08:00 | GitHub | noreply@github.com | 2025-04-03T18:32:54+03:00 | | vocab : use string_view::find() to avoid unnecessary looking up beyond the fragment range (#12706) |
| 2052 | 1c059995e080e37aeaacaae35850fc4c044b9dbe | 2004644b7a5da6fe080e51861ab583480280f1d3 | Jeff Bolz | jbolz@nvidia.com | 2025-04-03T10:08:26-05:00 | GitHub | noreply@github.com | 2025-04-03T10:08:26-05:00 | | vulkan: Fix missing cmake logic for dot product extension (#12721) |
| 2053 | 2004644b7a5da6fe080e51861ab583480280f1d3 | 5f696e88e0eb490c8028421d3c5419df9dcf5385 | Atharva Dubey | atharva.dubey@codeplay.com | 2025-04-03T13:12:39+01:00 | GitHub | noreply@github.com | 2025-04-03T15:12:39+03:00 | | ci : add env variable in ggml-ci and document the same in SYCL.md (#12736) |
| 2054 | 5f696e88e0eb490c8028421d3c5419df9dcf5385 | 193c3e03a63ccda3ac3d6a2999e41e6d1414fe23 | R0CKSTAR | xiaodong.ye@mthreads.com | 2025-04-03T19:51:35+08:00 | GitHub | noreply@github.com | 2025-04-03T13:51:35+02:00 | | sync : minja (inclusionAI/Ling) and update tests (#12699) |
| 2055 | 193c3e03a63ccda3ac3d6a2999e41e6d1414fe23 | 65cfe136a0793b2fdf5af8bdb1ab2cf10c3e1d2e | a3sh | 38979186+A3shTnT@users.noreply.github.com | 2025-04-03T15:32:55+08:00 | GitHub | noreply@github.com | 2025-04-03T09:32:55+02:00 | | fix MUSA compiler warning (#12704) |
| 2056 | 65cfe136a0793b2fdf5af8bdb1ab2cf10c3e1d2e | 3f9da22c2b21a2cef216de50006436ef1cab8764 | Chenguang Li | 757486878@qq.com | 2025-04-03T15:18:08+08:00 | GitHub | noreply@github.com | 2025-04-03T15:18:08+08:00 | | CANN: Support operator SIN COS ARGMAX (#12709) |
| 2057 | 3f9da22c2b21a2cef216de50006436ef1cab8764 | 2a0dc97e56eac6db0a4016f0b45da6d0a0055ef2 | Alan Gray | agray3@users.noreply.github.com | 2025-04-03T02:31:15+01:00 | GitHub | noreply@github.com | 2025-04-03T03:31:15+02:00 | | Simplify and improve CUDA graphs through use of indirect copy pointers (#9017) |
| 2058 | 2a0dc97e56eac6db0a4016f0b45da6d0a0055ef2 | 97a20c012be2f9bfbeee0209405a31001f93ccf9 | hipuddi |