- 61comments
- 749comments
- 973comments
- 154comments
- —discuss
- 387comments
- 191comments
- 21comments
- 49comments
- 455comments
- 138comments
- 398comments
- 7comments
- 141comments
- 101comments
- 94comments
- 347comments
- 15comments
- 107comments
- 9comments
- 222comments
- 34comments
- 49comments
- 6comments
- 12comments
- 111comments
- 209comments
- 28comments
- 13comments
- 536comments
Here you can see a 26B model hacked using a trained KV cache bank, responding like Gemma. The latency I am seeing is <127 ms.
That's awesome I came from the jev in python thread. is the model shareable? do you have any more writing on this? would love to read more about it! cool demo anyways.