Architecture questions

#2
by jbenjaminw - opened

A couple questions, as I'm currently working on my own model(s).

  1. Have you performed any ablation studies i.e. what are the advantages of your architecture vs. a naïve approach of just using a basic transformer?
  2. Why GDN vs. once again, the naïve approach of full attention? You're not dealing with a very large context window so the quadratic blowup at these scales likely isn't extreme and you may be sacrificing representational ability at the scales you're working with while adding computational complexity, reducing throughput and thus able to train on fewer tokens (and overtraining small models seems to be the go-to approach with memory limitations).

I ask because I think we're training on similar hardware and at similar scales and I'm frankly a bit bored by the idea of "attention is all you need" but paradoxically, the approach of "full transformer" may be the smartest choice at this scale as much as I don't like that answer. But I could be completely mistaken here; these are honest questions.

SurjoLabs org
  1. We did not do extensive ablation studies for this generation. However we did studies to ensure recurrent layers which is the backbone of our architecture is better when you are memory limited or parameter limited.
  2. Our current models are meant to be experiments for a larger model which will natively support large contexts. Using it here gave us some experience with GDN-2. Also GDN-2 does help during training even on lower context windows.

We did make some mistakes when designing the architecture for this model. We aim to fix them in Surjo2

What have you noticed that it does with smaller contexts? My understanding is that attention strictly has more representational capacity at the expensive of tracking a lot more things at once (hence the memory blowup)

SurjoLabs org

The cost of hybrid attentions are O(N) while not using it would be O(N²)
If we were to change the seq len of a training run from 1024 to 2048 normally we would have to square root the micro batch size to fit in memory. However in our runs we just have to halve it.

Sign up or log in to comment