- 13comments
- 73comments
- 758comments
- 981comments
- 175comments
- 389comments
- 194comments
- 51comments
- 465comments
- 23comments
- 408comments
- 144comments
- 144comments
- —discuss
- 9comments
- 40comments
- 95comments
- 102comments
- 348comments
- 110comments
- 22comments
- 232comments
- 9comments
- 14comments
- 37comments
- 55comments
- 115comments
- 209comments
- 30comments
- 1comments
Here you can see a 26B model hacked using a trained KV cache bank, responding like Gemma. The latency I am seeing is <127 ms.
That's awesome I came from the jev in python thread. is the model shareable? do you have any more writing on this? would love to read more about it! cool demo anyways.