Georgi Gerganov
113dd60005
server : bach has to be allocated for n_parallel sequences
2023-10-20 20:42:45 +03:00
FSSRepo
6b2437e32d
added thread safe pipeline
2023-10-20 12:07:32 -04:00
Georgi Gerganov
325d1793f7
server : minor sync
2023-10-19 15:03:24 +03:00
Georgi Gerganov
9740824ba5
server : snake case
2023-10-19 14:44:37 +03:00
Georgi Gerganov
e3a2c3fe32
server : use refs + use llama_batch_clear()
2023-10-19 14:44:04 +03:00
Georgi Gerganov
3d5929e8ee
server : bug fix in ingest_images
...
n_tokens is incremented internally by llama_batch_add
2023-10-19 14:43:19 +03:00
Georgi Gerganov
a8c981b734
server : remove beam-search functionality
2023-10-19 14:10:37 +03:00
Georgi Gerganov
654e0a1fe0
server : coding-style normalization (part 2)
2023-10-19 14:09:45 +03:00
Georgi Gerganov
e44ed60187
server : coding-style normalization
2023-10-19 13:50:23 +03:00
FSSRepo
ab2fc00224
latest changes of sampling API
2023-10-18 16:57:48 -04:00
FSSRepo
8540568c48
Merge branch 'master' of https://github.com/ggerganov/llama.cpp
2023-10-18 16:55:26 -04:00
FSSRepo
7196c4e08a
new sampling API
2023-10-18 16:50:09 -04:00
Georgi Gerganov
0e89203b51
speculative : add tree-based sampling example ( #3624 )
...
* sampling : one sequence per sampling context
ggml-ci
* speculative : add tree-based sampling support
ggml-ci
* speculative : reuse the n_parallel CLI param
* speculative : refactor sampling
* examples : fix build after sampling refactoring
ggml-ci
* batched : fix n_seq_id
* sampling : fix malloc
ggml-ci
* swift : fix build
ggml-ci
* swift : try to fix build
ggml-ci
* prompts : add assistant.txt
* common : add llama_batch_add() and llama_batch_clear() helpers
* speculative : minor refactor
ggml-ci
* minor : comments + rename
ggml-ci
* speculative : fix off-by-one for n_drafted
* speculative : fix the n_drafted fix + p constants
2023-10-18 16:21:57 +03:00
FSSRepo
c02c52efb5
fix multiple clients
2023-10-17 17:54:56 -04:00
FSSRepo
d2b1fac6c7
fix make bui;d errors
2023-10-17 17:18:56 -04:00
FSSRepo
ed0c11cb83
multimodal support enabled by default
2023-10-17 16:58:20 -04:00
FSSRepo
6c277eaab5
update api like OpenAI
2023-10-17 16:53:38 -04:00
FSSRepo
4d1804330e
fix llava implementation
2023-10-16 16:31:17 -04:00
FSSRepo
d7eca255d7
context shift fixed
2023-10-16 14:43:10 -04:00
FSSRepo
2d9f11db28
fixed premature end due stop word
2023-10-16 12:36:05 -04:00
FSSRepo
fd64f04fc2
fix long prompt than ctx proposed in #3639
2023-10-15 19:07:18 -04:00
FSSRepo
b727e022d6
fix ci make build undefined ref errors
2023-10-15 18:53:48 -04:00
FSSRepo
ce961a304b
some ci fixes
2023-10-15 18:46:01 -04:00
Steward Garcia
9035978aae
Merge pull request #6 from damian0815/fssrepo_mac_fixes
...
fix compilation errors with llvm
2023-10-15 18:38:52 -04:00
FSSRepo
4e5c5c451c
notify the user from server ui that multimodality is unavialable
2023-10-14 08:28:49 -04:00
Damian Stewart
299f6b54d8
fix compilation errors with llvm
2023-10-14 11:17:38 +02:00
FSSRepo
7e64bfe060
refactor code + remove unused comments + improved README.md
2023-10-14 00:31:34 -04:00
FSSRepo
9f72b44635
add multimodal input - alfa
2023-10-13 23:36:32 -04:00
FSSRepo
de35b47908
fixed tokens probs
2023-10-13 19:55:25 -04:00
FSSRepo
9d98cdda2c
llava multimodal integration
2023-10-13 18:42:44 -04:00
FSSRepo
eb08201227
add changes to README.md
2023-10-13 14:28:06 -04:00
FSSRepo
a2c2d98c16
add context swap
2023-10-13 14:12:50 -04:00
FSSRepo
b6d9e212e5
fixed timings per slot
2023-10-13 13:10:38 -04:00
FSSRepo
6358ae5f48
server ui now support multiple clients
2023-10-13 12:22:54 -04:00
FSSRepo
4ba5a5013d
chat.mjs support cached prompt + some fixes
2023-10-13 11:06:41 -04:00
FSSRepo
500ac7120e
cached prompt support
2023-10-12 21:16:12 -04:00
FSSRepo
83c2b3553a
grammar + no stream completion
2023-10-12 18:43:57 -04:00
FSSRepo
5b8e29de53
multiple client support
2023-10-12 17:09:12 -04:00
FSSRepo
81484805f0
completion endpoint working
2023-10-12 16:17:27 -04:00
FSSRepo
29c8cdd65d
refactored sampling function
2023-10-12 15:02:19 -04:00
FSSRepo
b716eeb72a
Merge branch 'master' of https://github.com/ggerganov/llama.cpp
2023-10-12 12:55:08 -04:00
FSSRepo
78504218b9
save dev progress
2023-10-12 12:51:48 -04:00
Georgi Gerganov
57dd55e2c7
server : fix kv cache management ( #3588 )
2023-10-12 09:29:04 +03:00
FSSRepo
471230202d
crash fixed
2023-10-11 19:48:15 -04:00
FSSRepo
63f99b1ea6
implementing parallel decoding in server example
2023-10-11 18:14:11 -04:00
Michael Coppola
a8bdd65525
server : add parameter -tb N, --threads-batch N ( #3584 )
...
Co-authored-by: Michael Coppola <info@michaeljcoppola.com>
2023-10-11 22:42:22 +03:00
Kerfuffle
70c29da118
common : fix mirostat state when using multiple sequences ( #3543 )
...
* Fix mirostat state when using multiple sequences
* Fix mirostat by completely refactoring sampling!
* Try to fix zig build.
* Export function to fetch/create default sampler states
Code formatting cleanups and add some comments
Silence a warning about id not being used when logging is disabled
* Apply some renaming suggestions.
Fix comments that were out of sync with the pull.
* Use more consistant naming convention for sampling contexts
2023-10-11 22:35:46 +03:00
vvhg1
11ea5c7d96
infill. : fix tokenization ( #3508 )
...
* infill tokens correction
* serverinfill tokens correction
* removing any leading whitespace from infill suffix and removing leeading space token from suffix when params.escape
* removing any leading whitespace from infill suffix and removing leeading space token from suffix when params.escape
* only rm when params.escape, rm space if possible which is added back or rm added space token
* only rm when params.escape, rm space if possible which is added back or rm added space token
* Revert "only rm when params.escape, rm space if possible which is added back or rm added space token"
This reverts commit 63ba0b621f
.
* fix interactive prompt escaping and fix server infill leading space handling
* rm unnecessary bool check
2023-10-10 10:31:21 +03:00
Jhen-Jie Hong
97af49fa39
server : reuse llama_sample_token common util ( #3494 )
...
* server : reuse llama_sample_token common function
* common : use n_probs for temperature sampling
2023-10-06 15:44:24 +03:00
Kenvix ⭐
45eba9369f
build : use std::make_tuple() for compatibility with older GCC versions ( #3488 )
2023-10-05 20:16:39 +03:00