Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression (https://zartbot.github.io)

128 points by mfiguiere 3 days ago | 10 comments | View on ycombinator

arikrahman 3 days ago |

I am very impressed with the KV Cache Compression work as well as the prefix cacheing making queries converge on practically free.

mmastrac 3 days ago |

I've been working with an automatic incremental context compactor enabled and it's been surprisingly helpful. It was particularly effective with DS41f - I think I was running at an effective session length of 5M, with the model running around 300k-400k and it was holding on both speed and intelligence.

TBH I also ran the 400tok/s preview and that was just nuts. I just let the thing compact over and over over the course of a day attacking a couple of tough problems

vivzkestrel 3 days ago |

404 on the blog page? https://zartbot.github.io/blog/

N_Lens 3 days ago |

[dead]

smy20011 3 days ago |

Removed