llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2025-01-11 19:21:46 +00:00

History

Georgi Gerganov b0f27361f3 sampling : avoid expensive softmax during greedy sampling (#9605 ) * sampling : avoid expensive softmax during greedy sampling ggml-ci * speculative : fix default RNG seed + set sparams.n_probs * Update tests/test-sampling.cpp Co-authored-by: slaren <slarengh@gmail.com> * sampling : add clarifying comment [no ci] --------- Co-authored-by: slaren <slarengh@gmail.com>		2024-09-24 09:03:17 +03:00
..
CMakeLists.txt	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00
llama-grammar.cpp	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-grammar.h	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-impl.h	common : reimplement logging (#9418 )	2024-09-15 20:46:12 +03:00
llama-sampling.cpp	sampling : avoid expensive softmax during greedy sampling (#9605 )	2024-09-24 09:03:17 +03:00
llama-sampling.h	llama : refactor samplers internal implementation (#9370 )	2024-09-08 15:52:07 +02:00
llama-vocab.cpp	llama : support RWKV v6 models (#8980 )	2024-09-01 17:38:17 +03:00
llama-vocab.h	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama.cpp	cuda: add q8_0->f32 cpy operation (#9571 )	2024-09-24 02:14:24 +02:00
unicode-data.cpp	Removes multiple newlines at the end of files that is breaking the editorconfig step of CI. (#8258 )	2024-07-02 12:18:10 -04:00
unicode-data.h	llama : reorganize source code + improve CMake (#8006 )	2024-06-26 18:33:02 +03:00
unicode.cpp	unicode : add <algorithm> (#9508 )	2024-09-17 09:51:15 +03:00
unicode.h	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00