Files

86 lines
15 KiB
Plaintext

ebce9e1a548c2329aa97f53b8b2ad50d968088dc e850ff4de25ad1221abc9bf3bb30df252ccd9c36 2026-05-27T09:38:00+08:00 Yunzhe Jia origin/runtime_replace Refactor cal-llm integration by removing runtime capacity loading and adjusting request handling
2716130010c58f1cdd9d94f0e2f8e047bc92dc66 b35bda124ccd4ed581dd6088d3ffa4931c2809b6 2026-05-18T08:40:54+08:00 Yunzhe Jia fix n_ctx_slot error by adding runtime capacity loading for cal-llm
1bec0db5a675ddc60fb793be2a746e8f3c8855dc 465df20c3e1f3f58d0f5531c796e0941a49c32b0 2026-04-30T16:02:59+08:00 wangbomeng origin/dev-ben bench: support 2 calbins
6787a53ab9b4897c5a7b51a4062d7b90e6697d35 e1a76abf66178ea85656479a5e19aea7785c09f4 2026-04-24T09:22:02+08:00 wangbomeng calrt:input of multimodel from fp32 to bf16
059009eeaf20a569abfbac5eb37dad3bb9376428 1b073661d1311b5300173993b9f31ee7241d8606 2026-03-10T07:04:22+08:00 wangbomeng calrt: add multimodel
ed8aa63320393512bdcfe4b05b5ae01ba91888e1 48bd26501b08a3f0bff1249db47f313641f7bebb 2025-11-03T18:01:59+01:00 Daniel Bevenius model-conversion : pass config to from_pretrained (#16963)
5a91109a5d7dab5d7adc40bedb397ede99a705b1 f8f071faddf32ea09f4234edb6e809b380a9ee26 2025-10-24T12:02:02+02:00 Daniel Bevenius model-conversion : add trust_remote_code for orig model run [no ci] (#16751)
4b9f4cb0f89a88de4bdf97727d0457b0c648804c 85e72271ba1ce78adf34fd8997803c991e617ca6 2025-09-23T13:59:34+08:00 Aaron Teo devops: add s390x containers (#15915)
37a23c17bdc9c99c9c6ad41168e4ced3724b72cd 138c87ce8bd558b2cc134ada7316a3dad8eb67ac 2025-09-22T14:13:51+02:00 Adrien Gallouët common : enable `--offline` mode without curl support (#16137)
28baac9c9f491c872e2c37762d3bd90446b005e9 1eeb523c3e0c7ffbd59469f5463dcbdecba3535e 2025-09-21T16:50:45+03:00 Georgi Gerganov ci : migrate ggml ci to self-hosted runners (#16116)
4ca088b036313c2d8e682f4cfeb7c29edd85d0b9 703f9e32c4eb3166f8d63007c26e31a1466c21af 2025-09-18T16:22:50+01:00 Eric Curtin Add resumable downloads for llama-server model loading (#15963)
28b5f190ef1dbea5edf82dbc8b4407b721fadd13 86587da03bd78df8f4e7d8b111a0c1d2494d6ed0 2025-09-10T15:29:12+08:00 Chenguang Li CANN: implement LRU cache for ACL graphs (#15814)
5d6688de08e73acc2532d668380801ed79d704eb 4fd1242bef6cb2325b4ff1c1a80f3b54b64508a6 2025-09-05T04:36:23+02:00 Daniel Bevenius model-conversion : add --embeddings flag to modelcard.template [no ci] (#15801)
46d9caa27a0281150e8cf082308c0f9e7576ebe5 5a0e3ef6f00c658fbae53797f02d5a360ebf8fec 2025-08-28T09:26:48+02:00 Daniel Bevenius model-conversion : add mmproj conversion target (#15628)
ef0144c087b33e5b8da42d529ac71aaf05cb49df 2721257e3e2c4c944ac8a08221113ee7cb503f1b 2025-08-05T04:29:25+10:00 Sam model: support GLM 4.5 family of models (#14939)
90083283ec254fa8d33897746dea229aee401b37 d4b91ea7b2da253e1355b503f0fcb7b428ce005d 2025-07-19T12:51:22-04:00 compilade imatrix : use GGUF to store importance matrices (#9400)
0aedae00e6fb48680324a5ac5da9cba0e35de6b5 6bdda13981d6c8189b7dc4f9fb8ecb91c21529f8 2025-07-10T18:20:13-06:00 Gabe Goodhart model : Granite Four (#13550)
4a5686da22057867c23bd4a6be941ddc8c51e585 98bab638fb28cf95a5a66dd2d51b40d6c8f6d69a 2025-07-09T14:59:57-04:00 compilade llama : support Jamba hybrid Transformer-Mamba models (#7531)
5364ae4ba53cc6367b8c8bf78876839122ca4e57 7c07ac244d59c833bf209582f6df019a77cdda59 2025-05-16T07:38:07-07:00 Diego Devesa llama : print hint when loading a model when no backends are loaded (#13589)
7f323a589f8684c0eb722e7309074cb5eac0c8b5 3eac209319a6726fd9687c6188fc6b916b65953d 2025-05-11T20:18:39+08:00 David Huang Add `--no-op-offload` to improve `-ot` pp perf in MoE models like llama4 400B (#13386)
0527771dd80bd18479dfaaa0a98be297fc3592bf 2189fd3b6327a1d17893694125da8edcf74a6468 2025-05-09T17:25:50+08:00 R0CKSTAR llama-run: add support for downloading models from ModelScope (#13370)
b2034c2b55b36b2192bdefb3b295db2a911370f5 06bb53ad9b6e6d92b6ab6979927530080b1c990c 2025-04-11T20:01:56+08:00 tastelikefeet contrib: support modelscope community (#12664)
02082f1519565fc7b49de211b28bc5404a69209b df4d20cd53d5bb6fc21c1dc65f026d53b566d097 2025-03-26T22:06:04+08:00 Ivy233 clip: Fix llama-llava-clip-quantize-cli quantization error under CUDA backend (#12566)
333820d7491cd31c707a340ff23b984a84e40154 c026ba3c23765a648ca27c7a15ecf179f8e27f26 2025-02-07T15:48:47+02:00 magicse llama : fix progress dots (#11730)
9dd7a0390feffcc1f4b17eb7692a6e43030d85af c0d4843225eed38903ea71ef302a02fa0b27f048 2025-02-06T13:41:37+02:00 Georgi Gerganov llama : add log about loading model tensors (#11699)
a5203b4465c5c87813936bde98170e25bb09024f df984e014714cba4c99ef894b20b51cbcef31b16 2025-01-27T17:42:09+04:00 lexasub llama : minor fixes for up llama load model speed (#11448)
c07e87f38bd0c22ec6dbc852ae50aaa1c64632d4 564804b79b78df1469ec8646869972de5e885ec4 2025-01-24T09:02:38+01:00 stduhpf server : (webui) put DeepSeek R1 CoT in a collapsible <details> element (#11364)
47182dd03fe04a4ffda5d7f4c8a109ae0056cf56 3e6e7a6bc2c4b980a0cf0fcb5cb3b79a965b5f14 2025-01-06T10:55:18+02:00 Georgi Gerganov llama : update llama_model API names (#11063)
4ddd199f6f6b980e0a7ed9f9b44efeae2fbdf5c4 a0974156f334acf8af5858d7ede5ab7d7490d415 2024-12-15T15:43:25-05:00 Bartowski llava : Allow locally downloaded models for QwenVL (#10833)
10bce0450f0c4d80087e06312b9dbbab3e87f16b 1f922254f0c984a8fb9fbaa0c390d7ffae49aedb 2024-11-25T19:30:06+01:00 Diego Devesa llama : accept a list of devices to use to offload a model (#10497)
feff4aa8461da7c432d144c11da4802e41fef3cf 0abc6a2c25272d5cf01384dda8ee8bfec4ba8745 2024-09-13T14:23:11+02:00 Xuan Son Nguyen server : add loading html page while model is loading (#9468)
67155ab7f5e47c01b62aa989eab30f517bf6dc67 5af118efdaf1098798a06b24fd8a557760e99631 2024-09-11T12:52:37+03:30 Farbod Bijary feat: Implements retrying logic for downloading models using --model-url flag (#9255)
54f376d0b92c6ff6feb1fa2ef8ed2022348100ba b2e89a327457179a34eae4d7de0d412ed945679c 2024-09-09T11:04:39+03:00 Radoslav Gerganov rpc : update README [no ci] (#9320)
8f1d81a0b6f50b9bad72db0b6fcd299ad9ecd48c a47667cff41f5a198eb791974e0afcc1cddd3229 2024-09-01T22:38:17+08:00 Molly Sophia llama : support RWKV v6 models (#8980)
84eb2f4fad28ceadd415a4e775320c983f4d9a7d 1262e7ed13ac197c944f15e1ddb083cb4f36cf65 2024-08-12T20:45:50+08:00 Frank Mai docs: introduce gpustack and gguf-parser (#8873)
86e7299ef5dff0f388922dc6fcbce009e99d8005 60d83a0149849e9217e4b8ae26e277a41aea906e 2024-07-06T15:32:04-05:00 Derrick T. Woolworth added support for Authorization Bearer tokens when downloading model (#8307)
0c7b3595b9e5ad2355818e259f06b0dc3f0065b3 7b2f4a7d193ef2475259bbe7656fcccfab4b1217 2024-06-15T18:53:40+02:00 Xuan Son Nguyen Add `cvector-generator` example (#7514)
57684331fc2d685f7d1f5775af0b9e47d1829833 b83bab15a5d2a1e7807d09613a9b34309d86cfaa 2024-05-24T18:14:42-07:00 Mikko Juola Make tokenize CLI tool have nicer command line arguments. (#6188)
b18532a4efeca8796fea8e36195c81cbfd596a4a fcda1128bc5f8eb7e1811708fe9d9867b9aec815 2024-05-22T16:10:46+02:00 slaren phi3 : duplicate rope factors in each layer (#7447)
b83cc3f5b303ff30c52874b2d5864dc6385ebf9f 9cb317f77e53067f7a138cc89ef7657148eae8e6 2024-05-11T09:46:09+02:00 Joan Fontanals llama : add Jina Embeddings architecture (#6826)
f98eb31c517c95960df1d0abc48002787f145f3b bc4bba364fb96d908f2698e908648df5e6f55e02 2024-05-08T18:16:38-04:00 compilade convert-hf : save memory with lazy evaluation (#7075)
4cc120c7443cf9dab898736f3c3b45dc8f14672b 24ee66ed0d908d156bd0d1747b63a636a495cd7a 2024-04-12T14:11:46+02:00 Daniel Bevenius infill : add download instructions for model (#6626)
f4183afe6a22f356ee222a710686ae7f83dbd949 b804b1ef77351d2a11be945462c6c251710476cb 2024-04-11T15:22:47+02:00 Daniel Bevenius scripts : add --outdir option to hf.sh (#6600)
4bcd6b959ca3991084ad1d8464caf2a734e29b1d 9b84ae1806cded4d6683c7b810925da5ead40607 2024-04-04T09:49:21+02:00 Daniel Bevenius common: remove duplicate check for curl (#6471)
08a0c0206075556e82aca0feafad530dcc5f1426 52604860f93063ef98863921da697576af1c7665 2024-04-03T15:07:05+02:00 slaren ggml : mul_mat_id use the same tensor for all the experts (#6387)
f482bb2e4920e544651fb832f2e0bcb4d2ff69ab 1997577d5e121568ae39f538021733ccd4278c23 2024-03-23T18:07:00+01:00 Pierrick Hymbert common: llama_load_model_from_url split support (#6192)
dba1af612926cbd4ebe2d876277af1e3305177e0 ee804f6223777019cf921e0d99cc24669313ab98 2024-03-22T19:00:01+01:00 Pierrick Hymbert llama_model_loader: support multiple split/shard GGUFs (#6187)
d01b3c4c32357567f3531d4e6ceffc5d23e87583 cd776c37c945bf58efc8fe44b370456680cb1b59 2024-03-17T19:12:37+01:00 Pierrick Hymbert common: llama_load_model_from_url using --model-url (#6098)
b5f4ae09c3244ae1644b67c03ed9f4227ab25ad2 dfbfdd60f90207404039c6578d709231496831d9 2024-03-16T16:46:29+01:00 Daniel Bevenius gritlm : add initial README.md (#6086)
c2101a2e909ac7c08976d414e64e96c90ee5fa9e 515f7d0d4fce41c752fc253acf30707c3be2531e 2024-03-08T17:31:00-05:00 compilade llama : support Mamba Selective State Space Models (#5328)
21b08674331e1ea1b599f17c5ca91f0ed173be31 6a87ac3a52668e117d97bcea07b529c93188b303 2024-03-05T16:08:35+08:00 Neo Zhang Jianyu [SYCL] fix mul_mat fault in CI/unit-test (#5862)
9731134296af3a6839cd682e51d9c2109a871de5 4a6e2d6142ab815c964924896891e9ab3e050632 2024-03-02T22:00:14+01:00 Pierrick Hymbert server: tests: passkey challenge / self-extend with context shift demo (#5832)
973053d8b0d04809836b3339a50f68d9c842de90 7c8bcc11dc61cf5930b70cd0168b84afcebe12a9 2024-02-22T00:42:09+01:00 slaren llama : fix loading models with shared tok_embd and output (#5651)
df845cc982e7e2ea7b9900e29d55b15338faa78d 6b48ed089377330cdb362970a51c1c89b6d857a8 2024-01-13T17:29:43+01:00 David Friehs llama : minimize size used for state save/load (#4820)
e7e4df031b9e29d4b55a4e0b0295187f6b213db1 584d674be622fbf1578694ada6e62eebedbfd377 2024-01-12T20:07:38+01:00 slaren llama : ggml-backend integration (#4766)
e790eef21ce659f5c16d59f8a5c8dcf6cde0692a 5537d9d36bfdb4379555431f574d3d78ce6e7955 2024-01-12T05:48:00-07:00 Zay llama.swiftui : update models layout (#4826)
eab67950068e4b125007d027232c47d2a5831cd0 d8d90aa343c22fe01429d3540e47ded87e9dcb9d 2024-01-11T12:41:39-05:00 Behnam M server : add `LOG_INFO` when model is successfully loaded (#4881)
7a9f75c38b5e62fe27b8a5a3ed823b4a3714024b 5c1980d8d4c4e0c0af77359f81cc44d90b3f250b 2024-01-11T02:12:05-05:00 Behnam M server : update readme to document the new `/health` endpoint (#4866)
cd108e641dbdedd8c5641c4cec1762f751f38136 57d016ba2d46a6e22517a31a75cebb48f9e234b6 2024-01-10T14:56:05-05:00 Behnam M server : add a `/health` endpoint (#4860)
57d016ba2d46a6e22517a31a75cebb48f9e234b6 329ff615699d32f596d4ebf8baba654c30064e0d 2024-01-11T01:09:53+11:00 Brian llama : add additional suffixes for model params (#4834)
3c0b585561d74a56977cf3a3844535ecc9e37972 e5804313a1edaf00726ed0b96ecced07accbf50c 2024-01-04T16:22:38+08:00 singularity llama.swiftui : support loading custom model from file picker (#4767)
441f51dca004debf8b275f1bdc08e0f1af7fd8f8 38b3de4658292582a8941a2be5c77b40ce6ac0f2 2023-12-29T19:23:27+09:00 Tamotsu Takahashi ci : build with CLBlast + ggml-opencl use GGML_API (whisper/1576)
d232aca5a73b290e218a2e48b91023d5e994203f 31f27758faf4a4bd08101a57c7ec3a473f771f86 2023-12-21T21:07:46+01:00 slaren llama : initial ggml-backend integration (#4520)
800a489e4a8be199122259a995b1ee9dd7fae320 f7f468a97dceec2f8fe8b1ed7a2091083446ebc7 2023-12-17T19:38:41+02:00 Georgi Gerganov llama.swiftui : add bench functionality (#4483)
bcc0eb4591bec5ec02fad3f2bdcb1b265052ea56 81bc9214a389362010f7a57f4cbc30e5f83a2d28 2023-12-07T13:03:17+02:00 Georgi Gerganov llama : per-layer KV cache + quantum K cache (#4309)
5aa365d88fdb8fdd430ef3fc141c7a5fd37c3502 52c8bc3cf312e1caf02d37bfb9d9d865cbe33594 2023-12-05T10:19:18-07:00 Kerfuffle llama : allow overriding GGUF metadata when loading model (#4092)
03562f3a86d6706eea9f4fc09b532946c191b34e 37c746d687d877bc11803e96b4dc5f378b83c0a0 2023-12-02T02:17:06+08:00 CausalLM llama : support attention bias on LLaMA architecture (#4283)
b18c66ca6eee4fe0465cff5042daf05005dc9ab2 f4d973cecb7368c985720ba9100ae6abba14806d 2023-11-30T22:43:08+01:00 Daniel Bevenius llama : fix alignment of general.name in print meta (#4254)
54b4df8886103b436a4bb3b60f4d84824f9e8868 46876d2a2c92e60579dc732cdb8cbd243b06f317 2023-11-06T23:43:59-08:00 Matthew Tejo Use params when loading models in llava-cli (#3976)
71e3718abdb2771b50c9606d3a7569623a0b0afe 238657db2364cfb728c694470a4a81702afea760 2023-11-01T08:04:02+02:00 Georgi Gerganov llama : refactor graph build code (#3837)
a5e7dbd6141128bfa3c40a19c2945a181df625d3 d3956aea53369455008159cc405ed4c496976692 2023-10-22T12:14:56-06:00 Kerfuffle llama : validate special token ids are in range when loading GGUF model (#3635)
0e76a8992c8200237bbc6471a53fb8796b3872f7 2db94d98eda56982d80238840b0652b4137a2a84 2023-09-28T20:40:11+02:00 xaedes train : finetune LORA (#2632)
ec893798b7a2a803466cc8f063051499ec3d96f7 45855b3f1c7bdd0320aa632334d0b3e8965c26c4 2023-09-28T19:04:36+03:00 Georgi Gerganov llama : custom attention mask + parallel decoding + no context swaps (#3228)
dc07dc492ef9640bbb82904d7c7679f7bdcf6d76 ad9ddcff6ef322db5cf13785bd7c856b610d242e 2023-08-30T02:25:50-06:00 Kerfuffle convert : various script cleanups/fixes + merges and special token handling (#2842)
44c117f41ee01c5ac8fb86bba041f08d8b87b46d 43033b7bb4858da4f591715b3babdf906c9b7cbc 2023-08-28T21:51:47+02:00 xaedes train : mem usage and other improvements (#2439)
95385241a91a616788a3bb76d12c9b7b2379ca2d 335acd2ffd7b04501c6d8773ab9fcee6e7bf8639 2023-08-23T20:33:05+01:00 Olivier Chafik examples : restore the functionality to import llama2.c models (#2685)
6381d4e110bd0ec02843a60bbeb8b6fc37a9ace9 dadbed99e65252d79f81101a392d0d6497b86caa 2023-08-21T23:07:43+03:00 Georgi Gerganov gguf : new file format with flexible meta data (beta) (#2398)
fff0e0eafe817eef429ecb64f892ab7bdae31846 417a85a0010519224cf154eb85d383ffeafeeead 2023-07-20T13:47:26+03:00 Georgi Gerganov llama : fix regression from #2000 - could not load no-mmap models
a17a2683d8fdb899ba497d0c28ccafb28c62efb6 31cfbb1013a482e89c72146e2063ac4362becae7 2023-07-06T09:17:50-07:00 tslmy alpaca.sh : update model file name (#2074)
e32089b2c20b1b87b22912f4a8b93fe01647d5b9 2347e45e7bdb09c9a7d74b2c0bc86c2b65f0c343 2023-06-13T21:04:40+02:00 xaedes train : improved training-from-scratch example (#1652)
8c0a10e64dbf60fd9946c0cd5e6f59690800b123 fa84c4b3e80199a5683438f062009c031a06c4fa 2023-06-12T14:31:36+03:00 Kawrakow metal : fix failure to load model (#1817)
ffb06a345e3a9e30d39aaa5b46a23201a74be6de 7552ac586380f202b75b18aa216ecfefbd438d94 2023-05-30T21:24:22+03:00 Henri Vasserman OpenLLaMA 3B support (#1588)
affc76edfdefa7b326f526e463cc65ff13fcfb92 ea600071cb005267e9e8f2629c1e406dd5fde083 2023-05-20T14:19:28+02:00 Johannes Gäßler cuda : loading models directly into VRAM, norm calculation on GPU, broadcasting for ggml_mul (#1483)
b9fd7eee57df101d4a3e3eabc9fd6c2cb13c9ca1 b608b55a3ea8e4760c617418538465449175bdb8 2023-05-12T00:23:08+03:00 Georgi Gerganov ggml : remove bit shuffling (#1405)
78ca9838ee36660a776e97e3391b6fb5dcaacf7f a017390358cdb23fffb30988dc84bb190d0403ca 2023-03-29T13:51:37-07:00 Justine Tunney Make loading weights 10-100x faster
563cdc391dde140f1084d1012234e8e6f57f881f 8d4a855c241ecb0f3ddc03447fe56002ebf27a37 2023-03-24T08:19:05-07:00 comex Support calling mlock() on loaded model data on Linux and macOS (#453)