MZ
The Null Hypothesis
← Back to articles

Deep Learning

A Deep Reading of L₀

From counting nonzero weights to learning gates, stochastic relaxations, and structured channel pruning.

Keywords: L0 regularization · neural network pruning · Hard-Concrete · stochastic gates · Gaussian gates · structured sparsity

An abstract neural network diagram with sparse connections, gates, and red and black data flow on a pale grid.
In this essay

A fresh breeze carries your hat away, but thankfully the railway-station guard catches it before it lands on the tracks. That happens just before your first day in a job that irritates most storybook heroes and their authors: you are a kind-hearted, conscientious train-ticket inspector.

For you, it does not matter how much a passenger is carrying, as long as they have a ticket—as long as they are present, or represented to you. Any character you meet who has no ticket is treated as invisible, as though they should not be there, and must get off at the next station.

That is our dear L0L_0 ticket inspector, too.

The L0L_0 inspector is always ready to count the weights in a neural network. It does not matter whether a weight is large or small: the inspector asks only whether it is there. By contrast, its distant cousin, L1L_1, measures how much each passenger is carrying and keeps adding up the weights.

But why count the weights that are present—or, mathematically, the nonzero weights—in a neural network? There are many reasons, but to begin with we are answering one question, or really two at once: how can we get a model with good predictions and good accuracy while using as few weights as possible? (Louizos et al., 2018; Oliveira et al., 2024)

That is the efficiency people talk about day and night.

Before you close the article because you are afraid of equations that will take two hours to look up—and then you will return having forgotten what you were reading—do not worry. This big beast is only our inspector in his official uniform, standing among his coworkers.

Play the ticket inspector’s round · 60 seconds

The conscientious inspector

θ =0.90-0.450.200.70-0.80
L₀weight contribution1010L₁weight magnitude−101
∥θ∥0=∑j1{θj≠0}=5\|\theta\|_0 = \sum_j \mathbf 1\{\theta_j\ne0\} = 5∥θ∥1=∑j∣θj∣=3.05\|\theta\|_1 = \sum_j |\theta_j| = 3.05

Figure 1: The L₁ sum changes with weight magnitude; the L₀ count changes only when a weight becomes exactly zero.

The L0L_0 “norm” counts the nonzero entries. For a vector of weights θ\theta,

∥θ∥0=∑j1{θj≠0}.\|\theta\|_0=\sum_j \mathbf 1\{\theta_j\ne 0\}.

Here, 1{θj≠0}\mathbf 1\{\theta_j\ne0\} contributes one when weight jj is nonzero and zero otherwise. The subscript jj is simply the index of a weight travelling through the neural network. This is the (L_0) definition used by Louizos et al. (2018).

For example, take this train of weights:

θ=(3,0,−2,0,5).\theta=(3,0,-2,0,5).

The inspector counts three nonzero weights, so ∥θ∥0=3\|\theta\|_0=3. If some weights grow and others shrink, what happens? The value of L1L_1 changes, but L0L_0 stays at 3, as long as none of the nonzero weights becomes exactly zero.

This is how Louizos and fellow researchers introduce L0L_0: it is a beautiful, elegant idea, but a rather strict one. The question L0L_0 asks is: should this weight stay on the train? But why count the nonzero weights?

Try it yourself: count what the inspector sees

python

The result is an (L0) count of 3 and an (L1) total of 10.

L0L_0 and neural-network compression

Imagine a neural network with one million weights. If all of them are nonzero, then

∥θ∥0=1,000,000.\|\theta\|_0=1{,}000{,}000.

At first, every weight appears to be on board. But after training the network, suppose we find that only 100,000 weights affect the result—whether the result is a classification or something else. Then

∥θ∥0=100,000.\|\theta\|_0=100{,}000.

We have reached a good prediction using fewer active parameters; weights are among the parameters. But how did we get there? The inspector is good at counting nonzero entries. So far, though, the inspector does not know who has—or deserves—a ticket. Here Louizos and colleagues enter the story again.

If you already know a little about neural networks, you might think: let us add L0L_0 to the training objective and let gradient descent—what I like to call “walking down the hill”—tell us what to keep.

I cannot explain neural-network learning in just two lines; I will give it more attention at the end of the series. For now, here is what we need. Training a neural network uses backpropagation: a way to work backward from the final decision toward the beginning. It is part of the process of walking downhill. Backpropagation tells us the direction of the slope, while the optimizer actually takes the step downhill.

Walking down the hill

We are now on top of the hill. I know there have been many landscape examples in this series, but I have put you on a train travelling through green fields. To get downhill, you first need to know which way the slope goes. That is the job of the derivative:

∂L∂w.\frac{\partial L}{\partial w}.

What happens if we change ww, the weight, by a small amount?

From the L0L_0 point of view, reducing a weight does not change the decision: the weight is still present. The optimizer looks at the weight and tries changing its value; then it looks for a change in L0L_0. But L0L_0 stays the same, so there is no gradual slope leading us to the foot of the hill—that is, to zero.

That is the inspector’s strictness. When a weight reaches zero, it does not gradually descend there: it jumps from a count of 1 to a count of 0.

The inspector asks only: is the weight present or not?

So the inspector cannot ask, “Can you gradually leave the train?” The inspector simply throws the passenger off.

Try it yourself: take a gradient step

python

This smooth example has a slope that points toward a lower loss; the hard (L_0) count does not.

The gate

Louizos and colleagues therefore separate what had been bundled into one number—the weight itself and the decision about whether it is needed—into two parts. This leads us to:

θj=θ~jzj.\theta_j=\tilde\theta_j z_j.

Here θ~j\tilde\theta_j is the original weight, before the keep-or-remove decision. The value zjz_j is the decision, or gate: it is open (1) if the weight stays and closed (0) if it is removed. The value θj\theta_j is the weight after it passes through the gate.

So:

original weight×participation decision=effective weight.\text{original weight}\times\text{participation decision} =\text{effective weight}.

Our inspector has developed. The inspector still watches for presence, but now there are two things: the weight a passenger carries and the passenger’s ticket (the weight’s contribution and its presence decision).

Now we have one question to answer: is the gate open?

This is the beginning of the solution to the learning problem. How do we decide when a gate opens or closes? How do we decide which gates receive 1 and which receive 0?

Since zz takes either 0 or 1, and those outcomes have probabilities, it makes sense to use a Bernoulli distribution:

zj∼Bernoulli⁡(πj).z_j\sim\operatorname{Bernoulli}(\pi_j).

If weight jj is useful, the probability that gate zjz_j is open should rise. If its contribution is small, that probability should fall.

Try it yourself: apply a gate

python

The closed middle gate makes its effective weight zero: [4, 0, 7].

Expected L0L_0: counting gates on average

The next step is to find the expectation of zjz_j. The symbol for expectation is E\mathbb E. The Bernoulli gate expectation and expected-count identity follow Louizos et al. (2018).

We have:

E∥θ∥0=∑jπj\boxed{\mathbb E\|\theta\|_0=\sum_j\pi_j}

This means:

The expected number of active weights is the sum of the probabilities that their gates are open.

Start with θj=θ~jzj\theta_j=\tilde\theta_jz_j, where θ~j\tilde\theta_j is the original weight and zjz_j is its gate. If we treat θ~j\tilde\theta_j as nonzero, then the only thing that determines whether the effective weight θj\theta_j is present is zjz_j.

If zj=1z_j=1, then θj=θ~j\theta_j=\tilde\theta_j: the weight is present and L0L_0 counts it as one. If zj=0z_j=0, then θj=0\theta_j=0: the weight is absent and L0L_0 does not count it.

Therefore, for one weight, its contribution to the L0L_0 count is simply zjz_j, because zjz_j is either 1 or 0.

With several weights, say z1,z2,z3z_1,z_2,z_3, the number of open gates in one particular draw is

∥θ∥0=∑jzj.\|\theta\|_0=\sum_j z_j.

For example, if the gates at one moment are (1,0,1,1,0)(1,0,1,1,0), then ∥θ∥0=3\|\theta\|_0=3, because three gates are open.

But zjz_j is not fixed during training; it is random:

zj∼Bernoulli⁡(πj).z_j\sim\operatorname{Bernoulli}(\pi_j).

This means zj=1z_j=1 with probability πj\pi_j, and zj=0z_j=0 with probability 1−πj1-\pi_j.

Now we can make sense of expectation. E[zj]\mathbb E[z_j] asks: if we repeat the experiment of opening and closing this gate many times, what is its average value? Since the gate is either 1 or 0,

E[zj]=1⋅πj+0⋅(1−πj)=πj.\mathbb E[z_j]=1\cdot\pi_j+0\cdot(1-\pi_j)=\pi_j.

For example, if πj=0.8\pi_j=0.8, this does not mean the gate itself has value 0.8. In each draw, the gate is still either 0 or 1. But across many draws we expect it to be open about 80% of the time, so its average approaches 0.8.

Now return to the number of weights:

∥θ∥0=∑jzj.\|\theta\|_0=\sum_j z_j.

Take the expectation. An important property says that the expectation of a sum is the sum of the expectations:

E∥θ∥0=E[∑jzj]=∑jE[zj]=∑jπj.\mathbb E\|\theta\|_0 =\mathbb E\left[\sum_j z_j\right] =\sum_j\mathbb E[z_j] =\sum_j\pi_j.

For a small example, suppose we have three gates with probabilities π1=0.9\pi_1=0.9, π2=0.2\pi_2=0.2, and π3=0.7\pi_3=0.7. Then

E∥θ∥0=0.9+0.2+0.7=1.8.\mathbb E\|\theta\|_0=0.9+0.2+0.7=1.8.

This does not mean there are 1.8 actual weights. In any one draw, the number of open weights is 0, 1, 2, or 3. But if we repeat the gate sampling many times, the average number of open weights approaches 1.8.

Put simply:

Each passenger has a probability of having a ticket.
Add up the probabilities of all the passengers having tickets, and you get the expected number of passengers on the train.

This is the important transition. Instead of trying to work directly with a hard count of zeros and ones, we now have πj\pi_j, a quantity that can change gradually—for example, 0.9, 0.8, 0.6, 0.3.

That leads straight to our next question:

If we have made the L0L_0 part easier this way, what is still difficult in the training objective itself?

Try it yourself: estimate the expected active count

python

The simulated average should be close to the theoretical expectation of 1.8.

The objective

Now the neural network must learn two things at once: to make good predictions and to keep as few components active as possible. This leaves us with the larger equation in our journey so far:

J(θ~,π)=Ez[L(θ~⊙z)]+λ∑jπj.\mathcal J(\tilde\theta,\pi) =\mathbb E_z[\mathcal L(\tilde\theta\odot z)] +\lambda\sum_j\pi_j.

Do not be afraid of it. It contains the two things we are trying to teach the network: make predictions and keep the model small. Let us take it apart.

The first part: prediction quality

Ez[L(θ~⊙z)]\mathbb E_z[\mathcal L(\tilde\theta\odot z)] is the prediction quality. We met something similar a few pages ago, but now, instead of taking the expectation of zz, we take the expectation of L\mathcal L, the prediction loss. The loss focuses on how bad the network’s prediction is. The symbol ⊙\odot means element-by-element multiplication.

Let the weights be (4,2,7)(4,2,7) and the gates be z=(1,0,1)z=(1,0,1). Then

(4,2,7)⊙(1,0,1)=(4,0,7).(4,2,7)\odot(1,0,1)=(4,0,7).

L(θ~⊙z)\mathcal L(\tilde\theta\odot z) means: calculate the prediction loss using the network after the sampled gates have switched off some of its weights.

The expectation is, roughly, an average. It answers: across the gate configurations produced by these probabilities, does the prediction remain good on average?

The second part: the inspector’s bill

The term λ∑jπj\lambda\sum_j\pi_j is the expected number of open gates, multiplied by λ\lambda. The parameter λ\lambda controls how much we charge for keeping gates open. A larger λ\lambda means fewer weights are allowed to remain; only contributions that matter enough will justify their cost. Each open gate becomes more expensive. A smaller λ\lambda gives prediction quality higher priority, while a larger λ\lambda gives weight removal higher priority.

During learning, the two sides of the equation argue. Prediction tries to open as many gates as needed to keep its quality high. Compression tries to close as many gates as possible. What settles the argument? The components—or gates—provide the evidence.

If a component matters, prediction loss wins: its probability πj\pi_j rises. If it does not matter, compression wins: its probability πj\pi_j falls.

We will need to think about the derivative, or gradient, of this equation. The term λ∑jπj\lambda\sum_j\pi_j is friendly to differentiation because πj\pi_j is continuous. The problem is Ez[L(⋅)]\mathbb E_z[\mathcal L(\cdot)]: we are still sampling zz from a Bernoulli distribution, which returns zero or one. That is the old problem again.

Try it yourself: change the sparsity price

python

This toy calculation shows how the penalty's contribution grows with λ\lambda; it is not a neural-network training result.

How does the neural network learn?

Remember backpropagation? It needs a chain, like a→b→ca\to b\to c, whose parts we can differentiate using the chain rule.

But here the intermediate variable zz follows a Bernoulli distribution. It is discrete: it is either 0 or 1. We cannot differentiate through that discrete sampling step in the ordinary way.

We have now reached Louizos’s next major step.

Step 7: Continuous relaxation: give the gate a middle ground

z = clip(s, 0, 1)s = 0.62z = 0.62
z = clip(s, 0, 1)10−0.5011.5s

Figure 2: Clipping preserves intermediate gate values while allowing exact zero and one.

So far, the gate has been:

zj∈{0,1}.z_j\in\{0,1\}.

That means 0 is closed and 1 is open. This jump is too abrupt for ordinary backpropagation.

Louizos therefore first introduces a new continuous variable:

sj∼q(sj∣ϕj)\boxed{s_j\sim q(s_j\mid\phi_j)}

and then defines the actual gate by clipping:

zj=clip⁡(sj,0,1).\boxed{z_j=\operatorname{clip}(s_j,0,1)}.

This is the general continuous-gate recipe (Louizos et al., 2018). Do not worry about qq or ϕj\phi_j yet. First, understand the idea.

Before, the gate could only be 0 or 1. Now we first create a continuous value sjs_j, which might be −0.7-0.7, 0.2, 0.65, 1.3, or any other real number. Then we pass it through clip⁡(sj,0,1)\operatorname{clip}(s_j,0,1).

What does “clip” mean?

Clipping means:

  • Any value below 0 becomes exactly 0.
  • Any value between 0 and 1 stays as it is.
  • Any value above 1 becomes exactly 1.

Mathematically,

clip⁡(s,0,1)={0,s≤0,s,0<s<1,1,s≥1.\operatorname{clip}(s,0,1)= \begin{cases} 0,&s\le0,\\ s,&0<s<1,\\ 1,&s\ge1. \end{cases}

This is the piecewise definition. For example, if sj=−0.4s_j=-0.4, then zj=0z_j=0. If sj=0.25s_j=0.25, then zj=0.25z_j=0.25. If sj=0.8s_j=0.8, then zj=0.8z_j=0.8. If sj=1.6s_j=1.6, then zj=1z_j=1.

The gate can now take values such as 0, 0.1, 0.4, 0.8, and 1, rather than only 0 and 1.

Why is this useful?

Because the gate can now move gradually. Instead of jumping from 1 to 0 in one sharp step, it can move through 1, 0.8, 0.6, 0.3, 0. This gives gradient-based optimization something much better to work with. The mountain climber sees a gradual slope instead of a sudden cliff.

But why use clipping at all?

This is important. You might ask: if we want continuity, why not use sjs_j directly?

Because we still want actual removal to be possible. If a gate were always a soft value such as 0.00001, it would technically still be nonzero. Remember that L0L_0 cares about zero versus nonzero. Louizos therefore wants both:

continuous behaviour that is easier to learn

and

exact zeros that produce sparsity

Clipping gives us both. If sj≤0s_j\le0, then zj=0z_j=0 exactly—not 0.00001, but a true zero. If sj≥1s_j\ge1, then zj=1z_j=1 exactly. Clipping creates genuine zeros and ones while allowing continuous values between them (Louizos et al., 2018).

The inspector’s analogy

The inspector no longer thinks in only two states. Before, there was a ticket or no ticket. During training now, there is a middle ground. We might have zj=0.8z_j=0.8, meaning the weight currently contributes at 80%, or zj=0.3z_j=0.3, meaning it contributes at 30%. If the underlying variable moves below zero, the inspector finally says, “That is it—you are removed,” and we get zj=0z_j=0. We now have a gradual path toward removal.

But notice a subtle point: the actual gate is zj=clip⁡(sj,0,1)z_j=\operatorname{clip}(s_j,0,1). The continuous variable before clipping is sjs_j. They are not the same. For example, sj=−0.7s_j=-0.7 gives zj=0z_j=0, while sj=1.4s_j=1.4 gives zj=1z_j=1. So sjs_j can range over a wider set of real values, while zjz_j is forced to remain within [0,1][0,1].

The probability of an active L0L_0 gate

The gate is active when zj>0z_j>0. Because of clipping, zj>0z_j>0 happens exactly when sj>0s_j>0. Therefore,

P(zj>0)=P(sj>0).P(z_j>0)=P(s_j>0).

If QjQ_j is the cumulative distribution function (CDF) of sjs_j, then P(sj>0)=1−Qj(0)P(s_j>0)=1-Q_j(0). Thus,

P(zj>0)=1−Qj(0).\boxed{P(z_j>0)=1-Q_j(0)}.

This is the continuous version of the Bernoulli idea. Instead of πj\pi_j, we now use 1−Qj(0)1-Q_j(0) as the probability that the gate is active.

The new objective becomes:

J(θ~,ϕ)=Es[L(θ~⊙clip⁡(s,0,1))]+λ∑j[1−Qj(0)].\mathcal J(\tilde\theta,\phi) =\mathbb E_s\left[\mathcal L\left(\tilde\theta\odot\operatorname{clip}(s,0,1)\right)\right] +\lambda\sum_j[1-Q_j(0)].

Do not memorize this equation yet. The important idea is:

binary gate ⟶ continuous random variable ⟶ clipped gate\boxed{\text{binary gate}\ \longrightarrow\ \text{continuous random variable}\ \longrightarrow\ \text{clipped gate}}

or, more briefly,

sj⟶zj=clip⁡(sj,0,1).\boxed{s_j\longrightarrow z_j=\operatorname{clip}(s_j,0,1)}.

These are the bridges that later let us use different continuous distributions. That is why this step matters when we reach Gaussian gates. Louizos’s framework here is broader than Hard-Concrete itself: the general idea allows other continuous distributions, as long as we can calculate the needed probabilities and use a suitable reparameterization.

This is where the road toward Gaussian gates begins. The next question is: how do we train the parameters of the distribution that generates sjs_j? This is where reparameterization enters.

Try it yourself: clip a continuous sample

python

Clipping maps values below zero to zero and values above one to one.

Reparameterizing randomness

At the last station, we saw that a continuous variable such as sjs_j is better than jumping directly between zero and one. But one important question remains: we are still sampling sjs_j from a probability distribution, so how can backpropagation work out how changing the distribution’s parameters changes the result?

Louizos uses an idea called reparameterization. The name is bigger than the idea. Instead of hiding randomness inside the sampling operation itself, we separate two things:

  1. A part we learn.
  2. A random part we do not learn.

In general, we write:

sj=f(ϕj,ϵj)\boxed{s_j=f(\phi_j,\epsilon_j)}

where ϕj\phi_j contains the distribution parameters we want to learn, and ϵj\epsilon_j is random noise drawn from a fixed distribution that does not depend on ϕj\phi_j.

The randomness has not disappeared. It has been moved into ϵj\epsilon_j; then an ordinary function connects it to the parameter we want to learn.

A quick example before Hard-Concrete

If the distribution is Gaussian, we can write:

ϵj∼N(0,1),sj=μj+σjϵj.\epsilon_j\sim\mathcal N(0,1),\qquad \boxed{s_j=\mu_j+\sigma_j\epsilon_j}.

Suppose at one step ϵj=0.6\epsilon_j=0.6, μj=0.4\mu_j=0.4, and σj=0.5\sigma_j=0.5. Then sj=0.4+0.5(0.6)=0.7s_j=0.4+0.5(0.6)=0.7.

During the backward pass, we treat the sampled value ϵj=0.6\epsilon_j=0.6 as fixed for that step. If we change μj\mu_j from 0.4 to 0.41, sjs_j moves from 0.7 to 0.71. There is now a clear path for differentiation to follow:

μj,σj⟶sj⟶zj⟶prediction⟶L.\mu_j,\sigma_j\longrightarrow s_j\longrightarrow z_j \longrightarrow\text{prediction}\longrightarrow\mathcal L.

Reparameterization does not remove randomness. It says: put the randomness in ϵ\epsilon, and make the parameter we want to learn visible inside an equation we can differentiate.

But be careful: this Gaussian example only explains the general idea. Louizos did not use a Gaussian gate in his main method. He chose Hard-Concrete, which we have finally reached.

Try it yourself: reuse fixed noise

python

The same noise is used each time; changing the learned location changes the sample smoothly.

Hard-Concrete: Louizos’s gate tries to bring both worlds together

We now want something that satisfies two demands that seemed to conflict: a soft gate that can be trained through, and exact zeros for some gates so we get real L0L_0 sparsity rather than merely small weights.

Hard-Concrete builds this in several steps, so do not try to take in the whole equation at once.

Step 1: a simple random number

We start with:

uj∼U(0,1).u_j\sim\mathcal U(0,1).

That means we draw a random number between zero and one, such as uj=0.2u_j=0.2 or uj=0.73u_j=0.73. There is no open-or-closed decision yet; this is only the source of randomness.

Step 2: turn it into Logistic noise

Louizos uses:

gj=log⁡uj−log⁡(1−uj).g_j=\log u_j-\log(1-u_j).

This transforms the number drawn from a Uniform distribution into Logistic noise. We do not need to dive into that distribution now. What matters is that gjg_j is the random part, and we will put something learnable next to it.

Step 3: introduce log⁡αj\log\alpha_j

Now:

qj=gj+log⁡αjβ,s~j=sigmoid⁡(qj).q_j=\frac{g_j+\log\alpha_j}{\beta}, \qquad \tilde s_j=\operatorname{sigmoid}(q_j).

Here is the first important quantity the network learns:

log⁡αj.\boxed{\log\alpha_j}.

Think of it as a handle that moves the gate distribution. Moving it in the positive direction makes the gate more inclined to open; moving it toward the negative direction makes it more inclined to close.

The parameter β\beta is the temperature. It controls how sharp or soft the transition is. You do not need to memorize the effect of every value yet. Just remember that log⁡αj\log\alpha_j is learned for each gate and β\beta controls the shape of the relaxation.

The sigmoid takes any real number and compresses it into (0,1)(0,1). So s~j\tilde s_j is now a soft value between zero and one.

But a value confined to (0,1)(0,1) cannot give us an exact zero

Here comes Hard-Concrete’s key move. Louizos does not stop at a value between zero and one; he stretches the range a little beyond both ends:

sˉj=s~j(ζ−γ)+γ\boxed{\bar s_j=\tilde s_j(\zeta-\gamma)+\gamma}

where γ<0\gamma<0 and ζ>1\zeta>1. The paper uses, for example, γ=−0.1\gamma=-0.1, ζ=1.1\zeta=1.1, and β=23\beta=\frac23.

Why stretch it? Imagine s~j\tilde s_j was trapped between zero and one. After stretching, some values can fall below zero and others can exceed one. Then we apply the move we already know:

zj=clip⁡(sˉj,0,1).\boxed{z_j=\operatorname{clip}(\bar s_j,0,1)}.

Values below zero become exact zeros, values above one become exact ones, and values between them remain soft.

Why is it called Hard-Concrete?

Concrete gives us a smooth continuous relaxation (Maddison et al., 2017). “Hard” refers to clipping, which creates an actual point mass at zero and at one.

The whole idea can be summarized as:

u⟶Logistic noise⟶sigmoid⟶stretch⟶clip.u\longrightarrow\text{Logistic noise}\longrightarrow\text{sigmoid} \longrightarrow\text{stretch}\longrightarrow\text{clip}.

In plain terms, the decision is no longer a ticket that suddenly appears. First there is a smooth path; at the end are two barriers. Anything falling below zero has the door closed completely, and anything passing one has the gate fully open.

The probability that a Hard-Concrete gate is open

We do not want only one sample zjz_j; we also need the probability that the gate is nonzero, so that we can calculate the L0L_0 penalty. For Hard-Concrete, this probability can be computed directly:

pj=P(zj>0)=sigmoid⁡(log⁡αj−βlog⁡−γζ)\boxed{ p_j=P(z_j>0)= \operatorname{sigmoid}\left(\log\alpha_j- \beta\log\frac{-\gamma}{\zeta}\right) }

Do not memorize it yet. The important point is that, as with πj\pi_j, we have a quantity pj=P(zj>0)p_j=P(z_j>0). We can add up ∑jpj\sum_jp_j to get the expected number of active gates.

When we minimize the objective, the L0L_0 penalty pushes these probabilities downward, while prediction loss pushes the probabilities upward for gates the model needs.

Does this mean every zjz_j is now zero or one?

No. During training a gate may be, for example, zj=0.37z_j=0.37 or zj=0.82z_j=0.82, because we are still using a continuous relaxation. At other times, clipping saturates it and makes it exactly zero or one.

So Hard-Concrete does not mean that every forward pass is a completely binary network. It is a clipped continuous distribution that can also produce genuine zeros and ones.

Louizos’s test-time gate is not the same as a training sample

This is important, and we will need it when considering practical use later. During training, we draw random samples. At test time, we do not want the model’s result to change every time a new random sample is drawn, so Louizos gives a deterministic gate based on the location it learned:

z^j=clip⁡(sigmoid⁡(log⁡αj)(ζ−γ)+γ,0,1)\boxed{ \hat z_j=\operatorname{clip}\left( \operatorname{sigmoid}(\log\alpha_j)(\zeta-\gamma)+\gamma, 0,1\right) }

Notice that it may still be fractional. A deterministic gate does not necessarily mean 0 or 1; it might be 0.6.

From now on, we must distinguish three things: the random training sample, the deterministic evaluation gate, and the final binary decision that actually removes something from the model.

Try it yourself: draw one Hard-Concrete gate

python

The last clipping step can produce an exact zero or one.

From one weight to a whole carriage: group sparsity

Output channelsC_out = 5
After compactionC_out = 5
c1
c2
c3
c4
c5
Retained5 / 5
Filter outputsC_out 5 → 5
Next-layer inputsC_in 5 → 5

Figure 4: A real speed benefit requires rebuilding the connected tensors after a channel is removed.

So far, we have spoken as if every weight had its own gate. But Louizos also allows an entire group to share one gate.

This brings us to the question we left open at the start: how can the method remove a whole channel instead of a single weight?

If we put one gate in front of a group of weights, zg=0z_g=0 closes the whole group, not just one weight. In a convolutional neural network, that group could be a complete output feature map.

Each channel has one gate, and that gate multiplies the entire feature map:

h~j=zjhj.\tilde h_j=z_jh_j.

If zj=0z_j=0, the whole feature map disappears—not one pixel and not one weight. The inspector’s ticket now belongs to an entire train carriage, not one passenger. If the carriage’s gate closes, everything inside it leaves the journey.

But what are we counting?

If we add only ∑gP(zg>0)\sum_gP(z_g>0), we are counting the expected number of active groups or channels. That is not necessarily the number of parameters, because different channels may contain different numbers of weights.

If group GgG_g contains ∣Gg∣|G_g| weights, then the expected number of active weights is:

E∥θ∥0=∑g∣Gg∣P(zg>0).\mathbb E\|\theta\|_0 =\sum_g |G_g|P(z_g>0).

So we must not confuse the number of channels, the number of weights, FLOPs, and actual speed. This distinction matters when we consider what pruning changes in practice.

Try it yourself: count expected groups and expected weights

python

The expected count is 1.05 groups but 5.8 weights because the group sizes differ.

What does a Louizos training step look like?

We can now see the whole journey in one training step.

First, we draw noise for the gates. Then we transform it through Hard-Concrete into zz. We multiply the weights by the gates:

θ=θ~⊙z.\theta=\tilde\theta\odot z.

Next, we run the forward pass and calculate prediction loss. Then we calculate the expected activity penalty from the gate probabilities:

λ∑jP(zj>0).\lambda\sum_jP(z_j>0).

We add the two terms. Backpropagation returns through the path we created with reparameterization and updates both the network weights and the gate parameters.

Louizos did not know in advance who deserved a ticket. Instead, prediction quality and the complexity penalty argue during training. Their interaction moves the gate parameters so that some parts become more likely to stay and others more likely to close.

Here we need to stop the train for a moment

After all this, it is easy to say: “So the Gaussian gate we will use next is Louizos’s method.” But that is not correct.

Louizos gave us the general framework for L0L_0 with clipped continuous gates, but his main method uses Hard-Concrete. The Gaussian stochastic gate has another direct source, which we will reach now: Yamada and colleagues.

From Hard-Concrete to Gaussian: the same structure, a different parent

The general idea we took from Louizos is:

continuous sample⟶clip⟶exact zero is possible\boxed{\text{continuous sample}\longrightarrow\text{clip}\longrightarrow\text{exact zero is possible}}

Alongside this, we calculate the probability that the gate is active from the CDF of the distribution that generated it.

Hard-Concrete uses Logistic noise and a sigmoid. Yamada says: we can use the same idea with a simpler Gaussian variable.

Yamada et al.: Gaussian stochastic gates

In Yamada’s method, each feature has its own stochastic gate (Yamada et al., 2020). The original expression can be written:

zd=clip⁡(md+ηd,0,1),ηd∼N(0,σ2).\boxed{z_d=\operatorname{clip}(m_d+\eta_d,0,1)}, \qquad \eta_d\sim\mathcal N(0,\sigma^2).

Using reparameterization, the same expression is:

ϵd∼N(0,1),sd=md+σϵd,zd=clip⁡(sd,0,1).\epsilon_d\sim\mathcal N(0,1),\qquad s_d=m_d+\sigma\epsilon_d,\qquad z_d=\operatorname{clip}(s_d,0,1).

Here mdm_d is the parameter we learn and σ\sigma is the amount of noise. In Yamada’s construction, the noise scale is fixed while we learn the location mdm_d.

What does mdm_d do?

Imagine a Gaussian curve moving along the number line. If we push mdm_d to the right, more of the distribution lies above zero, and the chance that the gate is open rises. If we push it to the left, more of the distribution moves below zero; after clipping, that part becomes zero.

So mdm_d is the handle that gradually decides whether the feature deserves to stay.

The activity probability for a Gaussian gate

s ∼ 𝒩(m, σ²)σ = 0.50P(z > 0) = Φ(m/σ) = 78.8%
−3−1.501.53s

Figure 5: This is the long-run probability of a nonzero gate, not the value of one sample.

The gate is active when zd>0z_d>0. Since clipping changes only values of sd≤0s_d\le0 into zero,

zd>0  ⟺  sd>0.z_d>0\iff s_d>0.

But sd=md+σϵds_d=m_d+\sigma\epsilon_d, so:

P(zd>0)=P(md+σϵd>0)=P(ϵd>−mdσ).P(z_d>0)=P(m_d+\sigma\epsilon_d>0) =P\left(\epsilon_d>-\frac{m_d}{\sigma}\right).

Using the symmetry of the normal distribution, we get:

P(zd>0)=Φ(mdσ),\boxed{P(z_d>0)=\Phi\left(\frac{m_d}{\sigma}\right)},

where Φ\Phi is the cumulative distribution function of the standard normal distribution. It tells us the area under the Gaussian curve to the left of a chosen value.

Example

If md=0.4m_d=0.4 and σ=0.5\sigma=0.5, then md/σ=0.8m_d/\sigma=0.8. Therefore,

P(zd>0)=Φ(0.8)≈0.79.P(z_d>0)=\Phi(0.8)\approx0.79.

This means the gate will be nonzero in about 79% of samples over the long run. Again, 0.79 is not the value of the gate itself. It is the probability that the gate is greater than zero.

The L0L_0 penalty in Yamada’s method

Instead of summing πj\pi_j, as we did with Bernoulli, or using the special Hard-Concrete formula, we now sum:

∑dΦ(mdσ).\boxed{\sum_d\Phi\left(\frac{m_d}{\sigma}\right)}.

This gives us the expected number of active gates. Once again, the objective has the form:

prediction loss+λ∑dΦ(mdσ).\text{prediction loss} +\lambda\sum_d\Phi\left(\frac{m_d}{\sigma}\right).

The L0L_0 penalty pushes mdm_d to the left to lower the probability of activity. Prediction loss pushes important gates in the opposite direction.

Try it yourself: compare the Gaussian probability with a simulation

python

The simulation should approach Φ(m/σ)\Phi(m/\sigma) as the number of draws grows.

Did Yamada remove channels?

No, not in the original paper. This is a point we must preserve.

Yamada mainly used the gate for feature selection at the network input: each input feature has its own gate. The original paper uses Gaussian gates for input-feature selection. How a related gate might support group-level choices is a question for a later part of this journey.

Why Gaussian instead of Hard-Concrete?

Yamada introduced Gaussian stochastic gates as an alternative to distributions built from Logistic noise, and reported more stable feature selection than Hard-Concrete in the experiments in that particular setting.

But we should not jump from that sentence to: “Gaussian is always better than Hard-Concrete.” That is not what the paper proves.

The careful statement is: they are two different ways to build the same general idea of a clipped continuous gate, and Yamada gave the Gaussian gate direct support in a feature-selection setting.

Hard-Concrete and Gaussian side by side

The shared structure is:

noise⟶continuous sample⟶clip⟶z∈[0,1].\text{noise}\longrightarrow\text{continuous sample} \longrightarrow\text{clip}\longrightarrow z\in[0,1].

With both methods, clipping can produce exact zeros, and we can calculate the activity probability and turn it into an L0L_0 penalty. What differs is the distribution that generates the sample and the formula for the activity probability.

Hard-Concrete:

P(zj>0)=sigmoid⁡(log⁡αj−βlog⁡−γζ).P(z_j>0)=\operatorname{sigmoid}\left(\log\alpha_j- \beta\log\frac{-\gamma}{\zeta}\right).

Gaussian:

P(zj>0)=Φ(mjσ).P(z_j>0)=\Phi\left(\frac{m_j}{\sigma}\right).

The inspector has not changed jobs. What changed is the machine that generates the probabilities of opening the gates.

A Hint of What’s to Come

A later part of this journey will ask what changes when one decision applies to a whole group rather than to one passenger at a time. We will consider when a learned gate becomes a practical structural choice, and why making a component inactive is not always the same as removing the computation it requires.

For now, keep this question in mind: when a whole group shares one decision, what should count as a ticket, and what does it mean for the group to leave the train?

Back to the inspector

At the beginning, our inspector knew only how to count passengers: present or absent. Then the inspector discovered that counting alone does not say who deserves a ticket.

We separated the weight from the decision about whether it exists. We made that decision probabilistic, then found a continuous path through reparameterization. Louizos built Hard-Concrete for that path; Yamada used a Gaussian route.

The central question remains:

How can we preserve good predictions while keeping fewer components—the ones that truly deserve to stay?

References

Louizos, C., Welling, M., & Kingma, D. P. (2018). Learning sparse neural networks through L0L_0 regularization. In Proceedings of the International Conference on Learning Representations. Full text.

Maddison, C. J., Mnih, A., & Teh, Y. W. (2017). The Concrete distribution: A continuous relaxation of discrete random variables. In Proceedings of the International Conference on Learning Representations. Paper.

Oliveira, F. D. R., Batista, E. L. O., & Seara, R. (2024). On the compression of neural networks using ℓ0\ell_0-norm regularization and weight pruning. Neural Networks, 171, 343–352. https://doi.org/10.1016/j.neunet.2023.12.019

Yamada, Y., Lindenbaum, O., Negahban, S., & Kluger, Y. (2020). Feature selection using stochastic gates. In Proceedings of the 37th International Conference on Machine Learning (Vol. 119, pp. 10648–10659). Proceedings and paper.

Read more