llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-12-27 20:04:35 +00:00

Author	SHA1	Message	Date
Georgi Gerganov	7494c78428	llama : sync gguf-llama with llama (#2613 ) * llama : sync gguf-llama with llama * tests : fix build + warnings (test-tokenizer-1 still fails) * tests : fix wstring_convert * convert : fix layer names * llama : sync gguf-llama.cpp * convert : update HF converter to new tokenizer voodoo magics	2023-08-14 21:33:33 +03:00
goerch	afc4ca2889	convert : update convert-new.py with tokenizer fixes (#2614 ) * Merge tokenizer fixes into the gguf branch. * Add test vocabularies * Adapt convert-new.py (and fix a clang-cl compiler error on windows)	2023-08-14 20:20:04 +03:00
goerch	ec1b100720	llama : tokenizer fixes (#2549 ) * Merge tokenizer fixes into the gguf branch. * Add test vocabularies	2023-08-14 19:30:28 +03:00
Georgi Gerganov	8af3a99ff1	Merge branch 'master' into gguf	2023-08-14 16:39:18 +03:00
Cheng Shao	d75561df20	server : add --numa support (#2524 )	2023-08-14 16:36:42 +03:00
Georgi Gerganov	f00780b2ee	llama : sync gguf-llama.cpp with latest llama.cpp (#2608 ) * llama : sync gguf-llama.cpp with latest llama.cpp * minor : indentation + assert * llama : refactor gguf_buffer and gguf_ctx_buffer * llama : minor	2023-08-14 16:28:44 +03:00
Georgi Gerganov	0c19ae70d5	simple : minor style changes	2023-08-14 12:58:12 +03:00
Jhen-Jie Hong	2feb8934eb	server : fix default grammar by use empty string in the UI (#2604 )	2023-08-14 16:20:17 +08:00
Jhen-Jie Hong	5517d6e692	server : implement json-schema-to-grammar.mjs & add grammar param in the UI (#2588 ) * server : implement json-schema-to-grammar.mjs by follow python impl * server : add grammar support in chat.mjs * server : implement grammer param in the UI * server : generate .hpp * server : remove trailing whitespaces * server : generate .hpp * server : fix sort of prop pairs * server : optimize regex & iteration	2023-08-14 15:16:54 +08:00
Georgi Gerganov	56a1f32072	Merge branch 'master' into gguf	2023-08-14 10:14:05 +03:00
M. Yusuf Sarıgöz	202eab04d3	gguf : quantization is working	2023-08-12 16:39:05 +03:00
M. Yusuf Sarıgöz	1fc3d30b71	gguf : start implementing quantization (WIP)	2023-08-12 16:09:47 +03:00
M. Yusuf Sarıgöz	b2571af255	gguf : start implementing quantization (WIP)	2023-08-12 14:28:17 +03:00
byte-6174	b19edd54d5	Adding support for llama2.c models (#2559 )	2023-08-12 01:17:25 +02:00
Equim	53dc399472	server: fixed wrong variable name in timing json (#2579 ) * server: fixed wrong variable name in timing json * remove redunct entry	2023-08-12 00:35:14 +02:00
DannyDaemonic	9ca4abed89	Handle `ENABLE_VIRTUAL_TERMINAL_PROCESSING` more gracefully on earlier versions of Windows.	2023-08-10 13:11:36 -07:00
Christian Demsar	e59fcb2bc1	Add --n-predict -2 for stopping generation on full context (#2565 )	2023-08-10 16:28:27 +02:00
M. Yusuf Sarıgöz	1c4d8bf981	gguf : start implementing libllama in GGUF (WIP)	2023-08-10 16:52:08 +03:00
Martin Krasser	1638757767	Fix grammar-based sampling issue in server (#2566 )	2023-08-10 13:16:38 +03:00
Martin Krasser	f5bfea0580	Allow passing grammar to completion endpoint (#2532 ) * Allow passing grammar to completion endpoint	2023-08-08 16:29:19 +03:00
chaihahaha	7ed8d1fe7f	llm.vim : multiline autocompletion, get rid of "^@" (#2543 )	2023-08-08 15:07:02 +03:00
Georgi Gerganov	e7f94d6fdc	vim : bring back simple llm.vim example	2023-08-08 15:06:18 +03:00
AustinMroz	2d7baaf50f	vim : streaming and more (#2495 ) * Update Vim plugin * Remove getbufoneline usage, Add input bind example. getbufoneline() appears to be a recently added function and has been replaced with getbufline for compatibility. An additional example that explains how to add a keybind that works in insert mode was added.	2023-08-08 14:44:48 +03:00
klosax	f3c3b4b167	Add --rope-scale parameter (#2544 ) * common.cpp : Add --rope-scale parameter * README.md : Add info about using linear rope scaling	2023-08-07 19:07:19 +02:00
Georgi Gerganov	8083ae347a	gguf : minor stuff	2023-08-07 19:02:18 +03:00
Georgi Gerganov	1da82c551f	Merge branch 'master' into gguf	2023-08-07 18:53:03 +03:00
DannyDaemonic	86c3219895	console : fix issue related to Windows 11 PowerShell console mode persistence (#2521 )	2023-08-06 09:49:34 +03:00
Jonas Wunderlich	332311234a	fix firefox autoscroll (#2519 )	2023-08-04 22:16:11 +02:00
Cebtenzzre	182af739c4	server: regenerate completion.js.hpp (#2515 )	2023-08-04 21:00:57 +02:00
DannyDaemonic	3498588e0f	Add --simple-io option for subprocesses and break out console.h and cpp (#1558 )	2023-08-04 08:20:12 -07:00
Stephen Nichols	5f631c2679	Fixing race condition in server and partial stream handling in frontend. (#2391 ) * Fixing race condition in server.cpp and partial stream handling in completion.js * Reverting assert edits. * Adding newline to eof	2023-08-04 13:37:24 +02:00
Borislav Stanimirov	ff966e7ca6	build : fix several cast and printf warnings (#2499 )	2023-08-04 13:07:21 +03:00
Evan Jones	8183159cf3	examples : generate JSON according to schema (#1887 ) * examples : add JSON schema grammars * complete JSON grammar * ensure primitive types can be used as root of schema * support integer type and adjust usage text	2023-08-02 22:05:44 -04:00
M. Yusuf Sarıgöz	cf365fbc20	gguf : gguf counterpart of llama-util.h	2023-08-02 11:13:56 +03:00
Eve	81844fbcfd	tests : Fix compilation warnings (Linux/GCC) (#2451 ) * fix hellaswag print format, cast away warning in test-double-float * c++11 cannot use designated initializers * add static to test-grad0.c internal functions * use memcpy in test-double-float.c * port c tests to c++ * use initializer list for ggml_init_params	2023-08-02 11:06:19 +03:00
Bono Lv	c574bddb36	fix a typo in examples/server/README.md (#2478 )	2023-08-01 14:54:28 +02:00
ebraminio	86aeb27734	server : Support dark mode (#2414 ) * server : Support dark mode So it respects user system light / dark settings. * Update index.html.hpp by running ./deps.sh	2023-08-01 10:56:23 +02:00
M. Yusuf Sarıgöz	bb42aefaeb	gguf : mmap tensor data example	2023-07-31 17:46:12 +03:00
Johannes Gäßler	0728c5a8b9	CUDA: mmq CLI option, fixed mmq build issues (#2453 )	2023-07-31 15:44:35 +02:00
M. Yusuf Sarıgöz	08dc8fd884	gguf : do not hardcode tensor names to read	2023-07-29 10:24:46 +03:00
klosax	3492f848d7	gguf : add gguf_find_key (#2438 ) * gguf.cpp : find key example * ggml.h : add gguf_find_key * ggml.c : add gguf_find_key	2023-07-28 23:45:24 +03:00
klosax	8a88e5855c	perplexity : add Hellaswag calculation (#2389 ) * common.h : add hellaswag / remove perplexity-lines * common.cpp : add hellaswag / remove perplexity-lines * perplexity.cpp : add hellswag scores / remove perplexity-lines * perplexity.cpp : clean up * common.h : change default param value * common.cpp : Change default param * perplexity.cpp : alter wording * common.h : alter wording * common.cpp : alter wording	2023-07-28 21:25:36 +03:00
Georgi Gerganov	d73b8d48b4	examples : fix whitespace	2023-07-28 21:05:08 +03:00
nhamanasu	34ae1caf7f	examples : server chat mode with llama2 (#2400 ) * add: server chat mode with llama2 * fix: remove the unnecessary last \n	2023-07-28 21:02:10 +03:00
Weird Constructor	d91f3f0c55	readme : fix the description of the Tail free sampling (TFS) method (#2431 )	2023-07-28 11:44:43 +03:00
Rand Xie	65cdf34bdc	llama : use n_embd_gqa instead of n_embd to handle llama-2 70B (#2433 )	2023-07-28 11:42:53 +03:00
Georgi Gerganov	d2b6ca13ad	gguf : add array support	2023-07-27 14:53:07 +03:00
Georgi Gerganov	d89533dff6	gguf : expose the gguf_type enum through the API for now	2023-07-27 11:10:34 +03:00
Georgi Gerganov	5628ec7163	gguf : read / write sample models	2023-07-26 22:40:45 +03:00
Georgi Gerganov	4d698495ea	gguf : init	2023-07-26 18:21:12 +03:00

1 2 3 4 5 ...

252 Commits