https://github.com/cilium/cilium eBPF-based Networking, Security, and Observability
🕸️🔗
🔧🔗https://github.com/IST-DASLab/marlin FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.