llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-12-26 11:24:35 +00:00

History

Georgi Gerganov cf658adc83 llm : add Falcon support (#2717 ) * llama : refactor GGUF constants into static maps * llama : check if model architecture is known * llama : refactor llama_model_load_internal() * gguf : add KV constant maps * llm : read arch-specific KVs * convert : add dummy scores + types * falcon : load tensor data (CPU only) * llama : fix loading progress bar * llama : add arch member to llama_model * falcon : CPU inference working * falcon : support non-40B models * falcon : minor * llama : minor updates ggml-ci * convert-falcon-hf-to-gguf.py : fix special token mapping * llama.cpp : llama default UNK token = id 0 * llama.cpp : fix bpe tokenizer * llama.cpp : fix the fix of bpe tokenizer * ggml : pass eps to ggml_norm * metal : implement RoPE (mode = 2) + avoid ggml_repeat * ggml : ggml_repeat always creates new tensor * falcon : copy-paste self-attention from LLaMA * metal : print extra compute pipeline info * falcon : minor changes (still chasing the Metal problem) * llama.cpp : fix linefeed token * metal : fix GELU kernel numerical stability by using precise::tanh * metal : temporary workaround for the concurrency optimization bug * falcon : add CUDA offloading (#2739) * llama : better model naming and size reporting * llama : prep new tokenizer support * llama : advanced BPE tokenizer based on ggllm.cpp imlpementation * llama : remove oboslete comment ggml-ci * common : remove obsolete BPE API + disable test-tokenizer-1 * llama : revert BPE special-case in llama_byte_to_token() * cuda : add TODOs for RoPE NeoX implementation * llama : default special tokens based on vocab type * perplexity : add log for start of tokenization --------- Co-authored-by: klosax <131523366+klosax@users.noreply.github.com> Co-authored-by: slaren <slarengh@gmail.com>		2023-08-23 23:08:04 +03:00
..
CMakeLists.txt	llm : add Falcon support (#2717 )	2023-08-23 23:08:04 +03:00
test-double-float.cpp	tests : Fix compilation warnings (Linux/GCC) (#2451 )	2023-08-02 11:06:19 +03:00
test-grad0.cpp	tests : Fix compilation warnings (Linux/GCC) (#2451 )	2023-08-02 11:06:19 +03:00
test-grammar-parser.cpp	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
test-llama-grammar.cpp	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
test-opt.cpp	tests : Fix compilation warnings (Linux/GCC) (#2451 )	2023-08-02 11:06:19 +03:00
test-quantize-fns.cpp	ggml : generalize `quantize_fns` for simpler FP16 handling (#1237 )	2023-07-05 19:13:06 +03:00
test-quantize-perf.cpp	ggml : generalize `quantize_fns` for simpler FP16 handling (#1237 )	2023-07-05 19:13:06 +03:00
test-sampling.cpp	ci : integrate with ggml-org/ci (#2250 )	2023-07-18 14:24:43 +03:00
test-tokenizer-0.cpp	llama : fix whitespace escaping in tokenizer (#2724 )	2023-08-23 00:10:42 +03:00
test-tokenizer-1.cpp	llm : add Falcon support (#2717 )	2023-08-23 23:08:04 +03:00