// Hacker Noon · 27 February 2026
dReLU Activation Function: Matching SwiGLU Performance with 90% Sparsity
Explore dReLU, a novel activation function that applies ReLU to both gate and up-projections. Achieve superior sparsity and lower validation perplexity without compromising model convergence or performance.
Hacker Noon
@hacker-noon · Language Models (dot tech)

hackernoon.com
Read Full Article at hackernoon.comHacker Noon@hacker-noon
Discussion 0
Loading
Got something to say?
or to join the conversation.