Random idea for a more generalized Matryoshka (tbf it's not really a Matryoshka anymore): why not just use weights for the individual dimensions?
rel_weights = math.log2(num_dims) - (torch.arange(num_dims) + 1).log2() + 1/2
This would result in the same general distribution of weights over the loss as the Matryoshka as we move toward the higher numbered dimensionsa, except smoothly, without the steps: each additional dimension adds less and less to the whole, so now it would then make sense to truncate the embedding to any arbitrary length, not just some randomly set values.