Efficient streaming language models with attention sinks 3 years agoThe authors just uploaded a FAQ section, which may clarify some of the confusions: https://github.com/mit-han-lab/streaming-llm/blob/main/READM... 0ThreadHN