llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-12-26 03:14:35 +00:00

History

Kawrakow 682986a08e Add Winogrande evaluation (#5015 ) * winogrande: simple implementation It doesn't look like it is working - why? For Mistral-7B it is barely better than random chance (score ~60% for 1267 tasks), while I see Mistral-7B scoring 78.4% on the HF leader board. 1-sigma statistical uncertainty for 1267 tasks is ~1.4, so no way the difference is due to statistics. * winogrande: somewhat better Score for Mistrali7-B is now 68.9 on the validation set of winogrande_debiased. Still far from the reported 78.4, but better than what I had before. * winogrande: improving Mistral-7B score is now 73.56. Still not quite 78.4 but getting there. We are also getting a lower score on HellaSwag compared to HF leader board, so I'm not expecting we will get up to 78.4 anyway. It looks like it is better to skip the choice word(s) when evaluating the average log-likelihood. This kind of makes sense because a more common word (in Winogrande this is often a name) will have a higher probability without knowing about the follow up context, and this will skew the log-likelihood towards the more common word. We can only do this if the choice words are not last in the sentence. It also looks like it is better to skip the punctuation at the end of the sentence, provided the choice words are not last. * winogrande: add dataset instructions --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>		2024-01-18 13:46:27 +02:00
..
base64.hpp	llava : expose as a shared library for downstream projects (#3613 )	2023-11-07 00:36:23 +03:00
build-info.cpp.in	build : link against build info instead of compiling against it (#3879 )	2023-11-02 08:50:16 +02:00
CMakeLists.txt	cmake : fix ld warning duplicate libraries libllama.a (#4671 )	2023-12-29 16:39:15 +02:00
common.cpp	Add Winogrande evaluation (#5015 )	2024-01-18 13:46:27 +02:00
common.h	Add Winogrande evaluation (#5015 )	2024-01-18 13:46:27 +02:00
console.cpp	check C++ code with -Wmissing-declarations (#3184 )	2023-09-15 15:38:27 -04:00
console.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
grammar-parser.cpp	grammar-parser : fix typo (#4318 )	2023-12-04 09:57:35 +02:00
grammar-parser.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
log.h	english : use `typos` to fix comments and logs (#4354 )	2023-12-12 11:53:36 +02:00
sampling.cpp	llama : apply classifier-free guidance to logits directly (#4951 )	2024-01-15 15:06:52 +02:00
sampling.h	llama : fix copy/paste error in llama_sampling_params comment (#4994 )	2024-01-17 09:17:50 +02:00
stb_image.h	examples: support LLaVA v1.5 (multimodal model) (#3436 )	2023-10-12 18:23:18 +03:00
train.cpp	train : fix typo in overlapping-samples help msg (#4758 )	2024-01-03 19:53:40 +02:00
train.h	sync : ggml (backend v2) (#3912 )	2023-11-13 14:16:23 +02:00