text-generation-webui

mirror of https://github.com/oobabooga/text-generation-webui.git synced 2026-03-18 11:24:39 +01:00

Author	SHA1	Message	Date
oobabooga	fb1b3b6ddf	API: Rewrite logprobs for OpenAI spec compliance across all backends - Rewrite logprobs output format to match the OpenAI specification for both chat completions and completions endpoints - Fix top_logprobs count being ignored for llama.cpp and ExLlamav3 backends in chat completions (always returned 1 instead of requested N) - Fix non-streaming responses only returning logprobs for the last token instead of all generated tokens (affects all HF-based loaders) - Fix logprobs returning null for non-streaming chat requests on HF loaders - Fix off-by-one returning one extra top alternative on HF loaders	2026-03-12 14:17:32 -03:00
oobabooga	387cf9d8df	Remove obsolete DeepSpeed inference code (2023 relic)	2026-03-04 17:20:34 -08:00
oobabooga	b5a6904c4a	Make --trust-remote-code immutable from the UI/API	2025-10-14 20:47:01 -07:00
stevenxdavis	dd6d2223a5	Changing transformers_loader.py to Match User Expectations for --bf16 and Flash Attention 2 (#7217 )	2025-09-17 16:39:04 -03:00
oobabooga	3b28dc1821	Don't pass torch_dtype to transformers loader, let it be autodetected	2025-08-05 11:35:53 -07:00
oobabooga	1d1b20bd77	Remove the --torch-compile option (it doesn't do anything currently)	2025-07-11 10:51:23 -07:00
oobabooga	b69f435311	Fix latest transformers being super slow	2025-07-09 19:56:50 -07:00
oobabooga	6c2bdda0f0	Transformers loader: replace `use_flash_attention_2`/`use_eager_attention` with a unified `attn_implementation` Closes #7107	2025-07-09 18:39:37 -07:00
oobabooga	d9de14d1f7	Restructure the repository (#6904 )	2025-04-26 08:56:54 -03:00
oobabooga	b3bf7a885d	Fix ExLlamaV2_HF and ExLlamaV3_HF after `ae02ffc605`	2025-04-20 11:32:48 -07:00
oobabooga	ae02ffc605	Refactor the transformers loader (#6859 )	2025-04-20 13:33:47 -03:00

11 commits