15 KiB
15 KiB
| 1 | ebce9e1a548c2329aa97f53b8b2ad50d968088dc | e850ff4de25ad1221abc9bf3bb30df252ccd9c36 | 2026-05-27T09:38:00+08:00 | Yunzhe Jia | origin/runtime_replace | Refactor cal-llm integration by removing runtime capacity loading and adjusting request handling |
|---|---|---|---|---|---|---|
| 2 | 2716130010c58f1cdd9d94f0e2f8e047bc92dc66 | b35bda124ccd4ed581dd6088d3ffa4931c2809b6 | 2026-05-18T08:40:54+08:00 | Yunzhe Jia | fix n_ctx_slot error by adding runtime capacity loading for cal-llm | |
| 3 | 1bec0db5a675ddc60fb793be2a746e8f3c8855dc | 465df20c3e1f3f58d0f5531c796e0941a49c32b0 | 2026-04-30T16:02:59+08:00 | wangbomeng | origin/dev-ben | bench: support 2 calbins |
| 4 | 6787a53ab9b4897c5a7b51a4062d7b90e6697d35 | e1a76abf66178ea85656479a5e19aea7785c09f4 | 2026-04-24T09:22:02+08:00 | wangbomeng | calrt:input of multimodel from fp32 to bf16 | |
| 5 | 059009eeaf20a569abfbac5eb37dad3bb9376428 | 1b073661d1311b5300173993b9f31ee7241d8606 | 2026-03-10T07:04:22+08:00 | wangbomeng | calrt: add multimodel | |
| 6 | ed8aa63320393512bdcfe4b05b5ae01ba91888e1 | 48bd26501b08a3f0bff1249db47f313641f7bebb | 2025-11-03T18:01:59+01:00 | Daniel Bevenius | model-conversion : pass config to from_pretrained (#16963) | |
| 7 | 5a91109a5d7dab5d7adc40bedb397ede99a705b1 | f8f071faddf32ea09f4234edb6e809b380a9ee26 | 2025-10-24T12:02:02+02:00 | Daniel Bevenius | model-conversion : add trust_remote_code for orig model run [no ci] (#16751) | |
| 8 | 4b9f4cb0f89a88de4bdf97727d0457b0c648804c | 85e72271ba1ce78adf34fd8997803c991e617ca6 | 2025-09-23T13:59:34+08:00 | Aaron Teo | devops: add s390x containers (#15915) | |
| 9 | 37a23c17bdc9c99c9c6ad41168e4ced3724b72cd | 138c87ce8bd558b2cc134ada7316a3dad8eb67ac | 2025-09-22T14:13:51+02:00 | Adrien Gallouët | common : enable `--offline` mode without curl support (#16137) | |
| 10 | 28baac9c9f491c872e2c37762d3bd90446b005e9 | 1eeb523c3e0c7ffbd59469f5463dcbdecba3535e | 2025-09-21T16:50:45+03:00 | Georgi Gerganov | ci : migrate ggml ci to self-hosted runners (#16116) | |
| 11 | 4ca088b036313c2d8e682f4cfeb7c29edd85d0b9 | 703f9e32c4eb3166f8d63007c26e31a1466c21af | 2025-09-18T16:22:50+01:00 | Eric Curtin | Add resumable downloads for llama-server model loading (#15963) | |
| 12 | 28b5f190ef1dbea5edf82dbc8b4407b721fadd13 | 86587da03bd78df8f4e7d8b111a0c1d2494d6ed0 | 2025-09-10T15:29:12+08:00 | Chenguang Li | CANN: implement LRU cache for ACL graphs (#15814) | |
| 13 | 5d6688de08e73acc2532d668380801ed79d704eb | 4fd1242bef6cb2325b4ff1c1a80f3b54b64508a6 | 2025-09-05T04:36:23+02:00 | Daniel Bevenius | model-conversion : add --embeddings flag to modelcard.template [no ci] (#15801) | |
| 14 | 46d9caa27a0281150e8cf082308c0f9e7576ebe5 | 5a0e3ef6f00c658fbae53797f02d5a360ebf8fec | 2025-08-28T09:26:48+02:00 | Daniel Bevenius | model-conversion : add mmproj conversion target (#15628) | |
| 15 | ef0144c087b33e5b8da42d529ac71aaf05cb49df | 2721257e3e2c4c944ac8a08221113ee7cb503f1b | 2025-08-05T04:29:25+10:00 | Sam | model: support GLM 4.5 family of models (#14939) | |
| 16 | 90083283ec254fa8d33897746dea229aee401b37 | d4b91ea7b2da253e1355b503f0fcb7b428ce005d | 2025-07-19T12:51:22-04:00 | compilade | imatrix : use GGUF to store importance matrices (#9400) | |
| 17 | 0aedae00e6fb48680324a5ac5da9cba0e35de6b5 | 6bdda13981d6c8189b7dc4f9fb8ecb91c21529f8 | 2025-07-10T18:20:13-06:00 | Gabe Goodhart | model : Granite Four (#13550) | |
| 18 | 4a5686da22057867c23bd4a6be941ddc8c51e585 | 98bab638fb28cf95a5a66dd2d51b40d6c8f6d69a | 2025-07-09T14:59:57-04:00 | compilade | llama : support Jamba hybrid Transformer-Mamba models (#7531) | |
| 19 | 5364ae4ba53cc6367b8c8bf78876839122ca4e57 | 7c07ac244d59c833bf209582f6df019a77cdda59 | 2025-05-16T07:38:07-07:00 | Diego Devesa | llama : print hint when loading a model when no backends are loaded (#13589) | |
| 20 | 7f323a589f8684c0eb722e7309074cb5eac0c8b5 | 3eac209319a6726fd9687c6188fc6b916b65953d | 2025-05-11T20:18:39+08:00 | David Huang | Add `--no-op-offload` to improve `-ot` pp perf in MoE models like llama4 400B (#13386) | |
| 21 | 0527771dd80bd18479dfaaa0a98be297fc3592bf | 2189fd3b6327a1d17893694125da8edcf74a6468 | 2025-05-09T17:25:50+08:00 | R0CKSTAR | llama-run: add support for downloading models from ModelScope (#13370) | |
| 22 | b2034c2b55b36b2192bdefb3b295db2a911370f5 | 06bb53ad9b6e6d92b6ab6979927530080b1c990c | 2025-04-11T20:01:56+08:00 | tastelikefeet | contrib: support modelscope community (#12664) | |
| 23 | 02082f1519565fc7b49de211b28bc5404a69209b | df4d20cd53d5bb6fc21c1dc65f026d53b566d097 | 2025-03-26T22:06:04+08:00 | Ivy233 | clip: Fix llama-llava-clip-quantize-cli quantization error under CUDA backend (#12566) | |
| 24 | 333820d7491cd31c707a340ff23b984a84e40154 | c026ba3c23765a648ca27c7a15ecf179f8e27f26 | 2025-02-07T15:48:47+02:00 | magicse | llama : fix progress dots (#11730) | |
| 25 | 9dd7a0390feffcc1f4b17eb7692a6e43030d85af | c0d4843225eed38903ea71ef302a02fa0b27f048 | 2025-02-06T13:41:37+02:00 | Georgi Gerganov | llama : add log about loading model tensors (#11699) | |
| 26 | a5203b4465c5c87813936bde98170e25bb09024f | df984e014714cba4c99ef894b20b51cbcef31b16 | 2025-01-27T17:42:09+04:00 | lexasub | llama : minor fixes for up llama load model speed (#11448) | |
| 27 | c07e87f38bd0c22ec6dbc852ae50aaa1c64632d4 | 564804b79b78df1469ec8646869972de5e885ec4 | 2025-01-24T09:02:38+01:00 | stduhpf | server : (webui) put DeepSeek R1 CoT in a collapsible <details> element (#11364) | |
| 28 | 47182dd03fe04a4ffda5d7f4c8a109ae0056cf56 | 3e6e7a6bc2c4b980a0cf0fcb5cb3b79a965b5f14 | 2025-01-06T10:55:18+02:00 | Georgi Gerganov | llama : update llama_model API names (#11063) | |
| 29 | 4ddd199f6f6b980e0a7ed9f9b44efeae2fbdf5c4 | a0974156f334acf8af5858d7ede5ab7d7490d415 | 2024-12-15T15:43:25-05:00 | Bartowski | llava : Allow locally downloaded models for QwenVL (#10833) | |
| 30 | 10bce0450f0c4d80087e06312b9dbbab3e87f16b | 1f922254f0c984a8fb9fbaa0c390d7ffae49aedb | 2024-11-25T19:30:06+01:00 | Diego Devesa | llama : accept a list of devices to use to offload a model (#10497) | |
| 31 | feff4aa8461da7c432d144c11da4802e41fef3cf | 0abc6a2c25272d5cf01384dda8ee8bfec4ba8745 | 2024-09-13T14:23:11+02:00 | Xuan Son Nguyen | server : add loading html page while model is loading (#9468) | |
| 32 | 67155ab7f5e47c01b62aa989eab30f517bf6dc67 | 5af118efdaf1098798a06b24fd8a557760e99631 | 2024-09-11T12:52:37+03:30 | Farbod Bijary | feat: Implements retrying logic for downloading models using --model-url flag (#9255) | |
| 33 | 54f376d0b92c6ff6feb1fa2ef8ed2022348100ba | b2e89a327457179a34eae4d7de0d412ed945679c | 2024-09-09T11:04:39+03:00 | Radoslav Gerganov | rpc : update README [no ci] (#9320) | |
| 34 | 8f1d81a0b6f50b9bad72db0b6fcd299ad9ecd48c | a47667cff41f5a198eb791974e0afcc1cddd3229 | 2024-09-01T22:38:17+08:00 | Molly Sophia | llama : support RWKV v6 models (#8980) | |
| 35 | 84eb2f4fad28ceadd415a4e775320c983f4d9a7d | 1262e7ed13ac197c944f15e1ddb083cb4f36cf65 | 2024-08-12T20:45:50+08:00 | Frank Mai | docs: introduce gpustack and gguf-parser (#8873) | |
| 36 | 86e7299ef5dff0f388922dc6fcbce009e99d8005 | 60d83a0149849e9217e4b8ae26e277a41aea906e | 2024-07-06T15:32:04-05:00 | Derrick T. Woolworth | added support for Authorization Bearer tokens when downloading model (#8307) | |
| 37 | 0c7b3595b9e5ad2355818e259f06b0dc3f0065b3 | 7b2f4a7d193ef2475259bbe7656fcccfab4b1217 | 2024-06-15T18:53:40+02:00 | Xuan Son Nguyen | Add `cvector-generator` example (#7514) | |
| 38 | 57684331fc2d685f7d1f5775af0b9e47d1829833 | b83bab15a5d2a1e7807d09613a9b34309d86cfaa | 2024-05-24T18:14:42-07:00 | Mikko Juola | Make tokenize CLI tool have nicer command line arguments. (#6188) | |
| 39 | b18532a4efeca8796fea8e36195c81cbfd596a4a | fcda1128bc5f8eb7e1811708fe9d9867b9aec815 | 2024-05-22T16:10:46+02:00 | slaren | phi3 : duplicate rope factors in each layer (#7447) | |
| 40 | b83cc3f5b303ff30c52874b2d5864dc6385ebf9f | 9cb317f77e53067f7a138cc89ef7657148eae8e6 | 2024-05-11T09:46:09+02:00 | Joan Fontanals | llama : add Jina Embeddings architecture (#6826) | |
| 41 | f98eb31c517c95960df1d0abc48002787f145f3b | bc4bba364fb96d908f2698e908648df5e6f55e02 | 2024-05-08T18:16:38-04:00 | compilade | convert-hf : save memory with lazy evaluation (#7075) | |
| 42 | 4cc120c7443cf9dab898736f3c3b45dc8f14672b | 24ee66ed0d908d156bd0d1747b63a636a495cd7a | 2024-04-12T14:11:46+02:00 | Daniel Bevenius | infill : add download instructions for model (#6626) | |
| 43 | f4183afe6a22f356ee222a710686ae7f83dbd949 | b804b1ef77351d2a11be945462c6c251710476cb | 2024-04-11T15:22:47+02:00 | Daniel Bevenius | scripts : add --outdir option to hf.sh (#6600) | |
| 44 | 4bcd6b959ca3991084ad1d8464caf2a734e29b1d | 9b84ae1806cded4d6683c7b810925da5ead40607 | 2024-04-04T09:49:21+02:00 | Daniel Bevenius | common: remove duplicate check for curl (#6471) | |
| 45 | 08a0c0206075556e82aca0feafad530dcc5f1426 | 52604860f93063ef98863921da697576af1c7665 | 2024-04-03T15:07:05+02:00 | slaren | ggml : mul_mat_id use the same tensor for all the experts (#6387) | |
| 46 | f482bb2e4920e544651fb832f2e0bcb4d2ff69ab | 1997577d5e121568ae39f538021733ccd4278c23 | 2024-03-23T18:07:00+01:00 | Pierrick Hymbert | common: llama_load_model_from_url split support (#6192) | |
| 47 | dba1af612926cbd4ebe2d876277af1e3305177e0 | ee804f6223777019cf921e0d99cc24669313ab98 | 2024-03-22T19:00:01+01:00 | Pierrick Hymbert | llama_model_loader: support multiple split/shard GGUFs (#6187) | |
| 48 | d01b3c4c32357567f3531d4e6ceffc5d23e87583 | cd776c37c945bf58efc8fe44b370456680cb1b59 | 2024-03-17T19:12:37+01:00 | Pierrick Hymbert | common: llama_load_model_from_url using --model-url (#6098) | |
| 49 | b5f4ae09c3244ae1644b67c03ed9f4227ab25ad2 | dfbfdd60f90207404039c6578d709231496831d9 | 2024-03-16T16:46:29+01:00 | Daniel Bevenius | gritlm : add initial README.md (#6086) | |
| 50 | c2101a2e909ac7c08976d414e64e96c90ee5fa9e | 515f7d0d4fce41c752fc253acf30707c3be2531e | 2024-03-08T17:31:00-05:00 | compilade | llama : support Mamba Selective State Space Models (#5328) | |
| 51 | 21b08674331e1ea1b599f17c5ca91f0ed173be31 | 6a87ac3a52668e117d97bcea07b529c93188b303 | 2024-03-05T16:08:35+08:00 | Neo Zhang Jianyu | [SYCL] fix mul_mat fault in CI/unit-test (#5862) | |
| 52 | 9731134296af3a6839cd682e51d9c2109a871de5 | 4a6e2d6142ab815c964924896891e9ab3e050632 | 2024-03-02T22:00:14+01:00 | Pierrick Hymbert | server: tests: passkey challenge / self-extend with context shift demo (#5832) | |
| 53 | 973053d8b0d04809836b3339a50f68d9c842de90 | 7c8bcc11dc61cf5930b70cd0168b84afcebe12a9 | 2024-02-22T00:42:09+01:00 | slaren | llama : fix loading models with shared tok_embd and output (#5651) | |
| 54 | df845cc982e7e2ea7b9900e29d55b15338faa78d | 6b48ed089377330cdb362970a51c1c89b6d857a8 | 2024-01-13T17:29:43+01:00 | David Friehs | llama : minimize size used for state save/load (#4820) | |
| 55 | e7e4df031b9e29d4b55a4e0b0295187f6b213db1 | 584d674be622fbf1578694ada6e62eebedbfd377 | 2024-01-12T20:07:38+01:00 | slaren | llama : ggml-backend integration (#4766) | |
| 56 | e790eef21ce659f5c16d59f8a5c8dcf6cde0692a | 5537d9d36bfdb4379555431f574d3d78ce6e7955 | 2024-01-12T05:48:00-07:00 | Zay | llama.swiftui : update models layout (#4826) | |
| 57 | eab67950068e4b125007d027232c47d2a5831cd0 | d8d90aa343c22fe01429d3540e47ded87e9dcb9d | 2024-01-11T12:41:39-05:00 | Behnam M | server : add `LOG_INFO` when model is successfully loaded (#4881) | |
| 58 | 7a9f75c38b5e62fe27b8a5a3ed823b4a3714024b | 5c1980d8d4c4e0c0af77359f81cc44d90b3f250b | 2024-01-11T02:12:05-05:00 | Behnam M | server : update readme to document the new `/health` endpoint (#4866) | |
| 59 | cd108e641dbdedd8c5641c4cec1762f751f38136 | 57d016ba2d46a6e22517a31a75cebb48f9e234b6 | 2024-01-10T14:56:05-05:00 | Behnam M | server : add a `/health` endpoint (#4860) | |
| 60 | 57d016ba2d46a6e22517a31a75cebb48f9e234b6 | 329ff615699d32f596d4ebf8baba654c30064e0d | 2024-01-11T01:09:53+11:00 | Brian | llama : add additional suffixes for model params (#4834) | |
| 61 | 3c0b585561d74a56977cf3a3844535ecc9e37972 | e5804313a1edaf00726ed0b96ecced07accbf50c | 2024-01-04T16:22:38+08:00 | singularity | llama.swiftui : support loading custom model from file picker (#4767) | |
| 62 | 441f51dca004debf8b275f1bdc08e0f1af7fd8f8 | 38b3de4658292582a8941a2be5c77b40ce6ac0f2 | 2023-12-29T19:23:27+09:00 | Tamotsu Takahashi | ci : build with CLBlast + ggml-opencl use GGML_API (whisper/1576) | |
| 63 | d232aca5a73b290e218a2e48b91023d5e994203f | 31f27758faf4a4bd08101a57c7ec3a473f771f86 | 2023-12-21T21:07:46+01:00 | slaren | llama : initial ggml-backend integration (#4520) | |
| 64 | 800a489e4a8be199122259a995b1ee9dd7fae320 | f7f468a97dceec2f8fe8b1ed7a2091083446ebc7 | 2023-12-17T19:38:41+02:00 | Georgi Gerganov | llama.swiftui : add bench functionality (#4483) | |
| 65 | bcc0eb4591bec5ec02fad3f2bdcb1b265052ea56 | 81bc9214a389362010f7a57f4cbc30e5f83a2d28 | 2023-12-07T13:03:17+02:00 | Georgi Gerganov | llama : per-layer KV cache + quantum K cache (#4309) | |
| 66 | 5aa365d88fdb8fdd430ef3fc141c7a5fd37c3502 | 52c8bc3cf312e1caf02d37bfb9d9d865cbe33594 | 2023-12-05T10:19:18-07:00 | Kerfuffle | llama : allow overriding GGUF metadata when loading model (#4092) | |
| 67 | 03562f3a86d6706eea9f4fc09b532946c191b34e | 37c746d687d877bc11803e96b4dc5f378b83c0a0 | 2023-12-02T02:17:06+08:00 | CausalLM | llama : support attention bias on LLaMA architecture (#4283) | |
| 68 | b18c66ca6eee4fe0465cff5042daf05005dc9ab2 | f4d973cecb7368c985720ba9100ae6abba14806d | 2023-11-30T22:43:08+01:00 | Daniel Bevenius | llama : fix alignment of general.name in print meta (#4254) | |
| 69 | 54b4df8886103b436a4bb3b60f4d84824f9e8868 | 46876d2a2c92e60579dc732cdb8cbd243b06f317 | 2023-11-06T23:43:59-08:00 | Matthew Tejo | Use params when loading models in llava-cli (#3976) | |
| 70 | 71e3718abdb2771b50c9606d3a7569623a0b0afe | 238657db2364cfb728c694470a4a81702afea760 | 2023-11-01T08:04:02+02:00 | Georgi Gerganov | llama : refactor graph build code (#3837) | |
| 71 | a5e7dbd6141128bfa3c40a19c2945a181df625d3 | d3956aea53369455008159cc405ed4c496976692 | 2023-10-22T12:14:56-06:00 | Kerfuffle | llama : validate special token ids are in range when loading GGUF model (#3635) | |
| 72 | 0e76a8992c8200237bbc6471a53fb8796b3872f7 | 2db94d98eda56982d80238840b0652b4137a2a84 | 2023-09-28T20:40:11+02:00 | xaedes | train : finetune LORA (#2632) | |
| 73 | ec893798b7a2a803466cc8f063051499ec3d96f7 | 45855b3f1c7bdd0320aa632334d0b3e8965c26c4 | 2023-09-28T19:04:36+03:00 | Georgi Gerganov | llama : custom attention mask + parallel decoding + no context swaps (#3228) | |
| 74 | dc07dc492ef9640bbb82904d7c7679f7bdcf6d76 | ad9ddcff6ef322db5cf13785bd7c856b610d242e | 2023-08-30T02:25:50-06:00 | Kerfuffle | convert : various script cleanups/fixes + merges and special token handling (#2842) | |
| 75 | 44c117f41ee01c5ac8fb86bba041f08d8b87b46d | 43033b7bb4858da4f591715b3babdf906c9b7cbc | 2023-08-28T21:51:47+02:00 | xaedes | train : mem usage and other improvements (#2439) | |
| 76 | 95385241a91a616788a3bb76d12c9b7b2379ca2d | 335acd2ffd7b04501c6d8773ab9fcee6e7bf8639 | 2023-08-23T20:33:05+01:00 | Olivier Chafik | examples : restore the functionality to import llama2.c models (#2685) | |
| 77 | 6381d4e110bd0ec02843a60bbeb8b6fc37a9ace9 | dadbed99e65252d79f81101a392d0d6497b86caa | 2023-08-21T23:07:43+03:00 | Georgi Gerganov | gguf : new file format with flexible meta data (beta) (#2398) | |
| 78 | fff0e0eafe817eef429ecb64f892ab7bdae31846 | 417a85a0010519224cf154eb85d383ffeafeeead | 2023-07-20T13:47:26+03:00 | Georgi Gerganov | llama : fix regression from #2000 - could not load no-mmap models | |
| 79 | a17a2683d8fdb899ba497d0c28ccafb28c62efb6 | 31cfbb1013a482e89c72146e2063ac4362becae7 | 2023-07-06T09:17:50-07:00 | tslmy | alpaca.sh : update model file name (#2074) | |
| 80 | e32089b2c20b1b87b22912f4a8b93fe01647d5b9 | 2347e45e7bdb09c9a7d74b2c0bc86c2b65f0c343 | 2023-06-13T21:04:40+02:00 | xaedes | train : improved training-from-scratch example (#1652) | |
| 81 | 8c0a10e64dbf60fd9946c0cd5e6f59690800b123 | fa84c4b3e80199a5683438f062009c031a06c4fa | 2023-06-12T14:31:36+03:00 | Kawrakow | metal : fix failure to load model (#1817) | |
| 82 | ffb06a345e3a9e30d39aaa5b46a23201a74be6de | 7552ac586380f202b75b18aa216ecfefbd438d94 | 2023-05-30T21:24:22+03:00 | Henri Vasserman | OpenLLaMA 3B support (#1588) | |
| 83 | affc76edfdefa7b326f526e463cc65ff13fcfb92 | ea600071cb005267e9e8f2629c1e406dd5fde083 | 2023-05-20T14:19:28+02:00 | Johannes Gäßler | cuda : loading models directly into VRAM, norm calculation on GPU, broadcasting for ggml_mul (#1483) | |
| 84 | b9fd7eee57df101d4a3e3eabc9fd6c2cb13c9ca1 | b608b55a3ea8e4760c617418538465449175bdb8 | 2023-05-12T00:23:08+03:00 | Georgi Gerganov | ggml : remove bit shuffling (#1405) | |
| 85 | 78ca9838ee36660a776e97e3391b6fb5dcaacf7f | a017390358cdb23fffb30988dc84bb190d0403ca | 2023-03-29T13:51:37-07:00 | Justine Tunney | Make loading weights 10-100x faster | |
| 86 | 563cdc391dde140f1084d1012234e8e6f57f881f | 8d4a855c241ecb0f3ddc03447fe56002ebf27a37 | 2023-03-24T08:19:05-07:00 | comex | Support calling mlock() on loaded model data on Linux and macOS (#453) |