Henkie Tenkie
HenkTenk
·
AI & ML interests
None yet
Organizations
None yet
Compatible with vLLM on Ampere?
➕ 4
7
#3 opened 2 months ago
by
HenkTenk
What is the update ?
9
#4 opened 3 months ago
by
robert1968
Garbage output
👍 2
6
#8 opened 3 months ago
by
andylele
PPLX or KLD, or other benchmark
1
#4 opened 3 months ago
by
HenkTenk
How does this compare to the original 8bit qwen quant and the 4 bit auto-round quant?
2
#5 opened 3 months ago
by
sparx3
Can't seem to correctly generate structured output
17
#3 opened 7 months ago
by
kldzj
Running on 6 GPUs
🤗 1
4
#10 opened 9 months ago
by
0xSero
KV Cache per token and Doubling the context size
3
#12 opened 6 months ago
by
HenkTenk
Unable to use on 4x3090
3
#4 opened 7 months ago
by
marutichintan
Running on 4 GPUs with TP=4
3
#11 opened 8 months ago
by
nephepritou
Duplicate files
➕ 1
7
#3 opened 8 months ago
by
darkstar3537
Improved useability
2
#7 opened 11 months ago
by
HenkTenk
Newest vLLM produces nonsense with longer context
4
#1 opened 11 months ago
by
HenkTenk
Does this actually work with VLLM?
32
#1 opened 12 months ago
by
sirus