Beyond Transformers: State Space Models as the Next Paradigm in AI - Badri Patro
About this talk
This talk explores the evolution of transformer models and their advancements, particularly in the fields of large language models (LLMs) and large vision models (LVMs). The speaker discusses applications such as question answering, image generation, and code writing, emphasizing the significance of generative AI architectures. He highlights current challenges such as privacy, bias in datasets, and issues stemming from AI's expansive capabilities, including hallucination phenomena where models produce false information. The necessity for effective fine-tuning methods is examined as a means to enhance the training process of LLMs, particularly for niche applications. Additionally, the speaker addresses various innovative strategies, including the integration of advanced architectures like the Mamba model and the use of lightweight models for specific tasks.
Full transcript
Today we'll discuss about something like transformer and beyond transformer, what kind of technology people are using building LLM and LVMs and all. >> [clears throat] >> So, I'm going through the different tech terminology. Wherever you feel you can ask questions. Uh Like basically, what is the use of that? And where we are using the LLMs and LVMs? Then the application and what sort of application we can
build up build on it. Then what kind of gen AI architecture you are looking for? And what's the current issue and what's you want to do the research and some direction like that. So, where you want to focus on it. So, if you see this this picture, so first tell me about PM's I ask this question to LLMs or language models. I got a answer that current
prime minister of India. So, I didn't tell that uh tell me about PM's of India. Just I told that tell me about the PM's. So, directly it took my MAC address, my address, everything, my personal details, my login information. Based upon that it find that location and it talk about the prime minister of India. It's not the problem of LLM. Also, if you go the large vision
model, can you generate a image of a uh current PM? So, it's also the same problem. So, it took you at the local information, the privacy information, whatever information it is there, it took some time to generate the image of the current look location. So, the third question, if you have a LVMs, large vision and language model, the the the second one is large vision model, and
the third one is large vision and language model. So, when you have a multimodal systems, also they're also saying kind of problems you will face. Just I I asked that what he's doing. So, based upon that, he's telling he's talking something like 10 important information about it. So, how will you solve this type of problems? And technologies used to reduce all bias and uh First of all,
the bias. And the other informationals, I'll give that when you have a large data, and where you your personal information are there, how will you do segregate training? Then how will you do the federated training type of system you you have to develop so that those kind of privacy information, those whatever the leaking, you can stop there. Hello. So, for that, let's say everybody knows about that
when I have a training data Uh so, I will get a generated similar to this my training data. So, this is where that diffusion model came to the picture, and the success of the diffusion model takes place, and it's looks like a realistic image like that. >> So, the basic application, like when you do that the transformer and the SSM based model, first of all, you can
do the question answering. The basic things are question answering, image generation, then music and video. And the fourth the fifth part is more most important and recently uh Sundar Pichai also noticed that 75% of the code is written by AI. So, whether do you need human or not, that is also a question. So, what sort of question you will ask uh uh like, can you design something
Riemannian field and which is working and not in uh Euclidean space and some different space, that also can code and give to you. You can just write uh can you merge two LLMs and build a new LLM uh mean two architecture and try to build a new architecture, it can write uh from start means data pre-processing, training training architecture, all kind of decision he can take and
try to design it. So, that's the beauty of the uh LLM and generating the code. Then the the problem is like uh recently you heard about that Amazon problem. So, Amazon what happened uh they try to fix a bug bug through the LLM and they update the code, the code is completely uh means written from the scratch, so that whatever whatever that production fixes are there, everything
is went out washed out. So, uh everything is washed out. Due to that uh a huge outage is came out in the imagine. So, you have to take care of that you you can't give everything to that agentic mode or LLM, so that generate code from scratch and they can do for uh means such type of production or outage kind of thing. The sixth part is the
data for the research. For a research, when you do the research in uh something like a medical domain and some uh like uh something like uh where you don't have a data, like, plane crash happened. So, you want to generate such type of things and try to prevent it. You don't have a data to train the model. So, how will you, uh, get a data? So, there
is something like a synthetic data genetic model. So, the, uh, last is the privacy preserving model. So, how will you train, uh, when you you have a information on your laptop or a PC, when you give a path from that you are talking as an take, uh, the data from, uh, some places. So, as an may explore some other, I mean, the different tools may explore. And
they can take your privacy data also and the model it LLM is trained on the privacy data and try to generate also. Means, whatever it's not there, that you can generate also. So, uh, the basic thing is the prompt based uh, generation method. Like, question answering, generating context, solving, uh, assessments and explaining the new concept. Like, I, uh, ask the Riemannian field, uh, that also I don't
know in uh, how will you explain that? You you have to go through the deep mathematics and you have to, uh, means, study something in that field. But, you ask such type of question to the LLM, they can give a simple answer that how, uh, you can also uh, tell that to give you some give some example of Riemannian field when the two models are merging. And
the image generation case, like, it is going for, as I told that image editing tool, Adobe is using that, uh, LLM based image editing tool, the logos, art generation. And try to generate a books and magazine. And the main problem is happening is right now the people are uh, you heard about that conference and send journals, right? So the most of the paper they are generating from
the LLM and they try to submit it a conference different conference. It is very difficult to identify who whether it is a written by the person or it is written by the large language models. So how will you detect such thing that is also one kind of research is going. Then music and video. So there is something like Saral. You heard about that model that can generate
a video within that can generate video for the two minutes. So if you increase your model capacity means increase your LLM tool the large like 570 billion parameter. So that can also build the instead of the two minute that can go for the five minute that can generate a picture for a one hour as well. So right now Microsoft also has released that model that email V2
version and Saral is also there. So you can generate the directly video for your paper whatever what sort of work you want to do you can generate a video for that. you heard about that this is a deep fake. So it's not Barack Obama is speaking. All right, let's give it a go. So yeah, audio is not there. So if you see in this problem is whatever
the they collect the lip sync model and they collect a text these are the Barack Obama is speaking that. So they just feed these two model these two data to the LLM model and they try to generate on it. Is it possible to anything at any point in time even if they would never say those things. for instance, they could have me say things like uh I
don't know, Killmonger was right. Or Ben Carson is in the sunken place. Or how about this, simply President Trump is a total and complete Now, you see, I would never say these things. At least not in a public address, but someone else would. Someone like Jordan Peele. This is a dangerous time. Moving forward, we need to be more vigilant with what we trust from the internet. Uh
it's a time when we need to rely on trusted news sources. May sound basic, but how we move forward in the age of information is going to be the difference between whether we survive or whether we become some kind of up dystopia. Thank you. And stay woke, Ben So, uh if you see, this is what happening right now. So, uh like people are if you know WhatsApp
University, lots more videos are like that it is coming. So, most of them they are the lip sync lip sync model is there. That just you need to put the script for that. So, also it is good. It I'm not saying that it's a bad. Like those who can't able to speak uh they want to present their paper and the research work in a conference, they just
let let's say uh the there are lots more application. So, I'm just talking about the one application. Like those who can't speak able to speak it. Also, those who are not want to Let's say I want to I want Japan and Tokyo and I want to present in their language. So, that is also so in that case. Like it's not a bad at all, something it is
bad also. Like, like this whatever you will see in the WhatsApp in your city, so those kind of thing try to avoid it. How will you avoid those kind of So, as I told that the code generation, the complete algorithm, everything is written from here. So, the extreme cases, like these are the extreme cases, like I want to train with a sunny day data, highway data, and
accident data, all kind of things happening. The edge cases, the harsh weather, those I can't able to get all the information. So, how will I collect those information? Then the privacy preserving model. So, if you have a data like a hospital day data, the real real sample data, how will you generate? Like, right now the third third works is there. First the platform and background you are
using that. So, what kind of authentic model you are using and all this based upon generative elements. So, they just gave a data source to the agent model, and they will tell that this is the LLM you are going to use. Your job is done. So, for the development side of you, it is very easy. For the policy designing policy, you have to design agent, you have
to design how this agent workflow will work. Then coming to that a fine-tuning the model. Fine-tuning SLM model. It's a very huge model, like LLMs are very huge. You can't able to train or a fine-tune from the scratch. So, you will come up with a packed base model, like LoRA, QLoRA, adapters. So, how will you fine-tune in it? Then you will tell, "What's the problem? I The
agent is there. Agent will take a data." But agent will take a data, but data is not The model had not seen the data. Just it's a feed the data and give it the output. If you train it so that the model has seen the data like okay, this is what input, this is what output I will need to the generate. So, then there may not be
the uh one problem is hallucination. Hal- hallucination problem is solved if you fine-tune the model. If you will not the fine-tune the model, as in can generate anything means baseless and whatever the things is that try to generate it, which is not useful that. To avoid this hallucination problem like uh uh I want to generate something like uh uh some table uh if you want uh something
like uh Microsoft cost for the different product. So, that can able to give without any fact, I'll say it can try to generate and uh just put a number on it. So, that type of hallucination let's say somebody has taken that data and try to present it. So, that person don't know because uh he's also there in the company, but they don't know about the real fact.
He just gave to the agent and agent is try to do that. In order to avoid that situation, you need to fine-tune it. So, the fine-tuning is not possible for the LLMs from the full fine-tuning. So, then you have to go for the adapter base So, uh again uh another uh incident to recapture like uh this is panda picture. Just you add a noise to that. So,
simple panda is there. Just you add a sim- small noise to this image. The output is given. You never see that uh the panda and given a a similar means there is no nothing like that the next word is coming or a given uh this the like the spaces is same they belongs to the same spaces nothing is like that but just adding a noise the models
can give a different results also. How to prevent such type of things? This is what the pixel space this is what the embedding space. I will go for that the basic transformer model when you have a transformer like this paper everybody heard about it attention all you need. So like you have a input you have a encoder you have a decoder. Then what's the problem with the
transformer people are not going for the transformer and you if you heard the recent meta and meta paper and also the GPT paper so they have something like the spectral models are there. So they are trying to avoid transformer because this multiplication. If you see this multiplication that cause a very huge. When you have big sentence now people are they don't know anything else so they will
give a simple PDF to the transformer and try to have a question answered and try to have it summary on it. So to do that transformer what it will take it will take all the words present in the PDF let's say there 20k or 30k words are there they try to generate a matrix they try to form this matrix. If you have a 20k or 30k so
20k cross 20k that much matrix is built. So you have to have a n square complexity on it. So your inference time is very slow to avoid that people are coming with the technology now the KV caching on it so when you're designing LLM and on it so you heard about the KV caching that the Q and V caching so they store the last few information where
wherever you ask the question similar to that question so they store in a memory but how much memory you want to keep it so how much your KV cache memory you want to keep it so that type of that is also one of the bottleneck to designing LLMs. So the LLMs are nothing but this model multiply with N become 1000. Where the transformer N become the 12
24 30 that that many layers are there when the layers are increased a huge like thousands or like that become a LLM. so this is what the transformer I'm not going to more detail on it like everybody heard about this transformer and this attention matrix this attention matrix is a main problem of the transformer people are try to avoid it so you have a matrix complexity of
N square how will you reduce to the N so if there are something like millions of tokens are there try to generate it they you are asking question answering question answering and you have a huge database you're feeding to the LLM so how will you solve that one so that is where the state space model is coming that complexity of N square become N so just I'm
giving the how transformer is working so each word is finding attention and each other word like here you can find how A is related here how robot is related so this is what the connection and finally if you see the number of uh here the attention, then you can see that uh how much the word is influencing with respect to that the other word and the nearest
word present in that. So, that many influence are attention you need to find it. So, then what kind of models are there? Like uh so, this is the problem in the transformers. And then what kind of model people are using for uh classification task like uh sentence classification, they want for the encoder only encoder. You don't need the encoder-decoder model. If you have a uh just you
try to generate a few sentence that is uh the decoder only. Then you have encoder-decoder base model. So, what sort of models we have developed like Roberta uh Cam Electra Mobile BERT and the Longformer. This is just a encoder base model where you want to do the classification task. This is what the Transformer-XL, XNet, GPT series, and the dialogue is uh Open GPT, OpenAI GPT is then
decoder base model. Then you have a encoder-decoder base model like T5, BERT. Uh these are the Then uh uh like as I mentioned, if you have a encoder model, then you can do for the name entity recognition task, then sentence classification task, then exact question answering task, uh extractive question answering task, masked language modeling task. The decoder is just based upon that text generation and the causal
language model. Then the encoder-decoder, if you want to do the summarization task, translation task, and generative question answering. So, uh if you just do the extractive question answering, it's like uh you just have a answer is that just you want to extract those answers present in the data set. So, these are the models are there. So, people are using uh based upon that application they are using
the transformer. So, next coming to the picture is that LLM. So, you know the LLM is the input is the tokenization is there, embedding is there, and the transformer is present. When you have a large context model, then how will you do the context embedding? Like the advanced patterning is there and the hidden process is there and the diffusion based uh architecture. After that, when you combine
these two, that is using the quantization. This is what that beauty is present on it. Embedding the sooner embedding and you can take a different embedding that is a different. But, the main concept is present by the diffusion and advanced patterning and hidden process on it. So, the uh large action model, so oh if you have a symbolic and the memory system and planning, but basically the
agentic based model, you want to uh oh have a generative agentic model. The large action model is based upon this. So, they they uh they have a task breakdowns, they have a planning, they have a memory, they have a neuro oh symbolic uh integration operations are there. So, after that they combining to the action execution, then they have a feedback. Then, the another beauty of the model
is coming to the mixture of expert model. So, oh So, this is uh where you want to use combining different task and they for the each expert is doing that one of the task. Finally, they you have a top case selection and they have a weighted output on it. Then, you have a uh vision large vision and language model. How this large vision and language model is
combining? You have a image input like a vision input, then you have a vision encoder like VIT and all. Then you have a text input and you have a text encoder. Then you have a projection interface, then after that you have a multi-processing, then the finally to combining that you have a large LLM model and give a output. Then next coming to this is that can I
use with the help of the small language model. I don't want a large language model. I I want to have a small language model. So they there the main difference is the contact compact tokenization and the quantized optimized embody. So the whatever that token tokens are used in LLM, if you want to compact it and you want to further specific application then you have first you have
to do that compact tokenization, then you optimize embedding process. That embedding is like LLMs, they have a a very huge dimensional embedding dimension that if you limited and if you fine-tune again that embedding model and finally you have a efficient transformer model. There are few mobile mobile nets are there and efficient transformers are there. F net is that one of the model people are using for efficient
transformer. Then after that they are doing the model optimization, then edge deployment. Can similarly that if you want the segmentation anything model. So here it is like I I want to segment like where is the projector in this room? Where is the projector is there? Where is the chairs are there? Means monitor is there. All kind of segmentation I want to have camera and all. So whatever
I want as a prompt that model will do and capture that segment that part and give that that kind of image. So, this is generally uh useful for the civil engineer where they are try to build a building and uh like apartments and buildings for the corporates that they are using like uh after that it is coming to tell where to use what kind of technology the
LLM versus rag versus authentic model. So, if you see LLM is a basic first user and a prompt that will give to the LLM and LLM will try to generate the answer. Where the rag is something like uh you have a database, you are having a embedding, you're feeding that embedding along with the prompt. So, uh here there uh LLM there is no database nothing like a
memory is there to store the data. So, here the context data is there and try to give a answer on it. Whereas the authentic model uh here along with the prompt context they have a memory, they have a plan, they have have execution, everything like uh coming and they give to the LLM along with the planning, they have a memory uh storage as well. They'll pull plan
and action. Then for based upon this thing along with the context and prompt that try to generate it. Authentic is very good, but uh you need to have something like the fine-tune on it. So, here you can just uh you're getting the data and all. If you fine-tune with the context, then the model is giving uh try to avoid such kind of uh hallucination problem on it.
So, this is what based upon the authentic uh model when you have a short memory, you have a long memory, then you have a what kind of the task and objective you are doing on it, then the what LLM is doing on it, what kind of action is taking based upon the planning, what kind of action is taking on it. The environment environment is like what sort
of the tools are there, the games, codes, simulation. It take take all these concepts and give a re-think store in the memory. And when the question you ask, when the query is asking, based upon that LLM is try to Uh so, whatever the box are there, here it is more detailed version of this book. So, this is what the complete authentic framework. If you have a multi-agent
system, like one agent is doing for the one task, like one agent is doing the vision-related task, when another agent is language-related task, you want to take a decision on it and work in it. It is a sequential flow where there is the different architectures are there. I'm not going through that. Sequential system and the parallel multi-agent system. So, you have a memory, you have a planning
agent, you have a LLM based upon that. So, all these the planning agent taking care of all this thing and feed to the LLM, then LLM based upon the user query, it try to give a answer on it. Like if you have let's say I have a all my YouTube videos based upon that I you want to ask some question on it. Like when I give a
talk on CVPR, so it will give the CVPR 2020, 2019, like that. So, that kind of question. So, one agent is collecting the videos, another is agent is extracting the information from the video and the third agent is looking for that where the CVPR is coming in that video. So, the multi-agent system is more useful when you have a different task and you have a follow questions
on it like when it is there in the Salt Lake when is there in the different place. So, that kind of things you can do that. So, it's a SN AI SN and agentic AI based system. So, before going to that the type of LM the LLM is used. So, if you see the GPT based the generative pre-trained transformer. So, where you have a the transformer architecture
along with that you have a This part is the MLP part. The MLP part is replaced by a here is that feed forward neural network or a MLP. Multi-layer perceptron is replaced by the mixture of expert model. Then you now right now it is coming with a large reasoning model uh instead of that retrieval task. So, people are not retrieving that they are trying to generate the
retrieval token. Then you have a LVMs like Queen LLM is there. But Queen V VAT is uh giving to the Queen LLM. Then similarly VAT is giving the Queen LLM. Finally, one more Queen LLM is there to combine it. So, similarly the architecture diagram of the SLM. So, how will you do that SLM and the action model and now the hierarchical language model uh it is also
coming place. So, what's the big one? The big one is the transformer So, just I will discuss about the challenge challenge occur and what kind of solution is there. So, if you see the the size of the data is used in 24 to 2023. So, data's the data size is from starting from 5 GB. It's in going and increasing increasing now the 10 TV is there. Now
now if you see that is a 13 terabytes of a data people are using they add all these things they use around the 13 terabyte of data to train the LLM model. So you never know what kind of tokens are how many tokens are required. If you see the massive task text generation that is a 2.34 tokenization when you want to purchase some LLM model so that
the price is based upon the how how much tokens you want to purchase means what kind of the tokenization model you want to take it. So if you have a huge tokenization your price will be more. So if you see that is a return uh uh the recent the sale price of all these AI based companies are going down because the ROI is very low. So the
return on investment is low because for each and every task so I want to classify something I don't require LLM in that case but now the people are using LLM to classify as They are doing a small small task through the LLM so each call is charging uh LLM is charging so the price is going very high. So whatever the companies are investing and they are not
getting the return on it. So that become a that become a ROI problem uh the we are facing right now. So how to solve this ROI so they come up with a some engineering task like heavy caching. So if somebody has asked in previously uh into that LLM they just try to give it they don't call to the LLM. So for each call you need to pay
you need to run your GPU whether your GPU cost is there for also your energy cost is there. How will you save such type of cost? the data the data size is going and increasing. So, how will you uh how will you have a clear data that is no duplicate present in the data? So, when you train on the model, so you can see that I don't
require such a huge data, so that I need to train on a cluster of GPUs. To avoid that, if you have something like duplicate in the data and the personal information in the data that you want to so break down it. Then the second part is the tokenization part. So, number of token necessary to convey the same you are making the 2.3 tokens. Can I Can I
reduce the SLMs? They They use something like a token optimization problem. So, just they have a some model on it. They give a input token and try to generate a token. Token to token trans learning is there. So, instead of you have a huge 2.3, you can have a small transformer model and present on it and try to generate a token net. So, those kind of things
small small improvement that did and they achieve they try to solve that problems in this is what the vocabulary and the tokenization problem can be improved in that. The token tokenization training cost is there. How to reduce this tokenization training cost? So, the high pre-training cost as I mentioned that you in order to avoid the halogenation, you need to train it. Or you you have a pre-trained
model and if you want to fine-tune it, how will you reduce pre-trained models like you take some pre-trained model model instead of fine-tune it to just use a adapter or a just use a Laura. Laura or a Q Laura. Q Laura is from Google. So, Laura or Q Laura or some adapt small adapter to When why why do you need fine-tune? Because when you have a new
data set, your own data needs that the new data set you want to give to the model, the model has to know your data. So, you just require to fine-tune it. For the small data, you just use the adapter base method to do >> [cough] >> So, there is like Arduino-ing and SD-noising method to avoid that. Then the fine-tuning overhead. So, you want to fine-tune the LLM
1 LLM 2 You can't use the same model to the fine-tuning all these things like sentiment model. If you want to fine-tune the sentiment analysis task, that can't be used for that question QA model, that can't be used a hate speech When you do LLM and try to use all these tasks, the thing is like transformer you can do it, but LLM you have to bear the
fine-tuning on it. So, why this cost? Because it's it's like 40 billion or 130 140 billion parameters are there for Lama model I'm talking about. If you go further to Google model, Google has something like 540 billion of parameter. So, when you load those model, you need to have a 800 or a 800 GPUs on it. To run 800 per hour, it is $100. So, to train
on a such a huge 13 TB of data, you can't imagine how much cost is involved on it. Uh So, building LLM or a fine-tuning also it's a huge cost. At least you need to uh fine-tune for 1 or 2 days. So, you uh cost you have to bear for your fine-tuning model. So, that is the occur in case of uh uh fine-tuning. To solve that one,
there is a PEFT-based method like uh uh VPT, the visual prompt fine-tuning. For the language case, it is a PEFT-based method. Adapters are also belongs to the PEFT. You use uh any uh the prefix tuning tuning and the prompt tuning and uh the LLM adapter or a multimodal adapter uh or you can take uh library PEFT library from Hugging Face. Just you fine-tune it to save the
model. So, these are the other model that is the long low-rank adaptation, MeZO, I I I A3, QoRA. Uh then then the the the another point is high inference latency. So, any input you have to give that the LLM will make a matrix on it. So, that matrix is cost that inference time. Uh to compute that matrix, that much time is required. So, that is what the
inference cost is required. So, how will you reduce inference cost? Cost, one solution is the the uh the solution is a So, KV cache is there. You can use quantization-based method. Instead of uh you are using like 32-bit floating-point operation, instead of that you can use a 4-bit. Uh yeah, your uh uh re re means answer regulation is bit less, but uh your if you want the
inference time that medical application and all where the inference time is latency is also matter. For that little accuracy is not accuracy is also matter, but you need to decrease like 32 to 8-bit. So, eight time reduction happen in the inference time. Another way for doing is instead of passing through the complete network, use a mixture of expert network. Only one expert will activate and that can
do it. The software challenge is also to like Megatron is Nvidia and Microsoft If these are the LLMs, then deep DeepSpeed is another DeepSeek is another company. DeepSpeed is another company where you want to go go train your LLM very fast. DeepSpeed technique you can use. So, there there are lots more method software companies are there. They want to do the efficient training on LLM model. Then,
limited context the limited context is another problem the backbone of the LLMs are there the transformer. So, what people are telling they they just the complete they want to generate a complete sentence, they break part by part and they feed it and try to generate it. Because the long range dependency is not there in the transformer. Because transformer each he has to calculate. Let's say I have
a four or five words, the sixth word I want to generate, complete matrix I need to compute because all this matrix multiplication the rows and column that must again I need to calculate it. The values are changing. That's the problem of the long range dependency. If your token length is more, the transformer model is not doing good to perform on it. To solve that one, there is
something like control system based theory that is a state space model. The complete act as a state space and These are the prompt bottleneck and This is for the hallucination problem. If you don't uh fine-tune it, then your model is leading to just if you go for the agentic way solution, your model lead to the uh have a fire hallucination problem. Also, you want to take set
a temperature and minimize the temperature and try to do hallucination, but the temperature set means you're you're reproducing again and again the same thing, but it is not guaranteed that whatever is generated, that is based upon the fact. The the hallucination Another kind of hallucination is This time I ask the same questions like uh what sort of LLM techniques and uh the pep techniques are there to
fine-tune LLM. So, first time you will get a one answer. Second time same question you will ask, you can't reproduce exact sentence. You will uh that will come a different sentence. So, to avoid that, they are uh minimizing the temperature. You heard about the temperature scaling problem. So, uh temperature is one kind of having a re- setting the reproducibility. But, hallucination is something like that that whatever
I told that uh you want to generate some tables and uh based upon your data or you summarize that table, you have a big uh uh I will show there are more lots more big table than whatever I have in my PPT. So, are you if you want to summarize, so what it will give it will just take some numbers and it will feed put in a
summarization. So, that kind of hallucination, how will you avoid that? that LLM text generator like uh it's a problem when you writing some articles and all. It's completely generated by LL LLMs. So, that may lead to like a problem that that is not based upon the fact fact and that is not read by the human. Just it's a machine paper. So, that may be harmful for uh
medical research and all. So, there are n number of the problem where you want to do the research. These are the research direction. If you want do a company and each and every problem is based upon like some fact and you want to develop and solve it. How will you solve it in efficient manner? I'm just talking about this limited context problem. So, next few slides I
will discuss about the limited context problem. How will you solve this limited context problem? Like long range dependency. So, attention models and the LLM models, they are not good on a long range dependency. How will you solve that? To do that, this is what the control system mechanism based method. You heard about gated MLP. You heard about H3. So, there is something like the Mamba. The linear
time sequence modeling with selected space. So, there is something that recurrent neural network people are trying. So, that is very fast and they have a long very fast and they can have a problem of vanishing gradient and they can't use a pre-training work. Then the people are coming with a transformer model where they want to have a pre-trained base model, so they can have a pre-trained model
and they solve the vanishing gradient problem. But the transformer problem has itself one more problem is that they try they don't have a long-range dependency. Uh where is that if you see HT minus one and HT, so these are uh the sequential model. The transformer model also fell facing kind of uh the forgetting or they don't have a dependency on it. How will you solve that combining
this RNN and the transformer? So there is something like if you heard of you know uh our plus two and like that the digital electronic there is a state space. One state is there, I want to go to the another state. How will I go from state one to state two? What kind of means uh like token one to the token two or token three. I want
to go from any state to any state. What is that cost? So how will I build a model so that it with a linear complexity I will go from state one to the state any state present on it. So uh this is what the uh the basic backbone uh is there. This is a simplified Mamba base architecture for vision. So people are using the self-attention network. remove
that spectral mixing. They remove that attention with spectral mixing. Spectral mixing has a problem. The complexity is L log N. Still it is not uh linear complexity. Then they they come up with the MLP base mixing method. MLP base is uh like uh they don't have a structural dependencies present on it. Then there is the one famous model is known as the Mamba. Mamba is a very
famous model the Nvidia is recently did with a 2 billion parameter. They tried to beat the steer all the model LM models with just a 2 billion parameter they based upon the Mamba 2 architecture that is called then they have a hyena, they have a hippo, then they have then we proposed that there is something like instability present Mamba when you want to work in a vision
based model. So instability what is instability when you train on a large network so any end problem is occurring the gradient instability because it's a control system based theory so they want to have something like exponential e to the power minus one minus delta so that e to the power of delta and the log term cause the in Mamba. So Mamba solves most of the things but
it has a instability problem occurs when you train on a large network like thousand layers and all. So to solve that one we come up with a method is in FFT Einstein fast Fourier transform method. So it's basically the signal system based theory when when you know that when the your system is unstable then you can have a unstable you should stay means open coming excited that
you should have a negative eigen value. Just you store those negative eigen value and remove all other eigen values present on that. That fast Fourier transform. So finding eigen value is not a simple in case of a huge parameter. So to do that we come up with a method is the fast Fourier transform. When you have a fast Fourier transform so the FFT is a efficient way
of finding eigen value. And after that we are just keeping that negative eigen values and store on it. Then another problem occur when you do that in FFT the parameter what should be that model parameter it So, to solve that what we come up with the Einstein matrix multiplication. It's not a simple matrix multiplication where you melt multiply these two matrix. Let's say you have a matrix
of dimension n cross b. N is that number of that tokens and the b is that the embedding dimension b or a d is the embedding dimension. You required a same n cross b matrix multiplications are there. How will you reduce that matrix multiplication? That is also a cost because that ultimately remove your hardware cost or a inference To do that if you see just the embedding
dimension is splitted and just multiply in a efficient in a embedding dimension. So, finally it will give a efficient matrix multiplication on it. So, this is this paper is accepted in ICASSP and due to this Einstein multiplication and that eigen value or theory. So, if you see this is what that standard matrix multiplication that is that stand tensor blending method where after that you can see there
is something like Einstein blending method if you required a very small number of the parameter on it. uh this is what we evaluated in the standard time series network when you have a seven time series data and you have a these are the models and leads to when I want to summarize this table. See, my model is the Simba TS is this. So, MSC is leading that
and here it is MSC and MA is doing very good. If I want to summarize, see if there is huge number of the text. If you summarize this table, LLM usually fails to do that because they just to give you some value to do that. So, how will you avoid that? That is also one question on We not only use the time series, we use for that
vision task like given a image segregate in thousand category. The problems I have discussed like the computational complexity that is quadratic with sequence length. Then you have a memory, you have a computation cost and the latency. To solve that, we develop another method is known as a SBT, the scattering vision transformer, which use a spectral network. It's a parameterized spectral network. Since you know that spectral transformation,
they are not parameterized, they are just transform from physical space to the spectral domain. Then we use Einstein matrix blending both in channel direction and the sequence direction. This is what that network we have in our network architecture that leads to have a scattering layer. The scattering layer is the scattering transform. They have a low frequency component, high frequency component, they have a inverse scattering transform. This
is what the parameterized one. The the spectral domain they had to don't have a parameterized form. When you do the parameterized form, what are the frequency I want to pass from the layer one to the layer two? What are the frequency are the required in each and every layer? The first layer is edge detection like that. The second layer is the contour, third layer is the object
level. So, what are the things are there? Then another poison there in that Mamba is the scanning method. If you have that along with this standard Mamba D has a unidirectional scanning, bidirectional scanning, like cross bidirectional the cross scanning. So, this is recent CVPR we published a year we are going to present this in CVPR as well. So, we remove this time uh the scanning method that
is known as the Hansa method. If you see the Hansa spectral network is much better than that V Mamba and vision Mamba paper. So, this is what the Hansa Hansa architecture is there. They have a state the the Hansa state space model, which is more efficient way of for mixing. They have a gated modulation, they have a spectral pulse net, and with the help of the Sagu
is a another gating architecture we used on it. So, this is much more efficient without scanning with to reduce the instability present in the Mamba architecture. So, based upon that we achieve a state state of art model, and it is 85.4 85.7 with very less if you see the floating point operation and the parameter as compared to the other model. So, we not only train on a
standard data set, also we validate with different data set how this model is working. It's working much better compared to the existing. Hansa is used can be used for the trans transport learning kind of network. Not only the transform learning or the task learning also. If you train on a one task and test on a different task, that can also so This is for the object detection
and instance segmentation task. Then finally the performance so evaluated for something on like parameters the flops, latency, throughputs, memory, and energy. All we combine and the we present this metric when you have something like LLMs or SLM, you have to evaluate like this what's the throughput, what's the memory consumption, what is the energy consumption, what is the latency. So, what kind of architecture it is there. So,
based upon that we are achieving much better than the existing one. Yeah, somewhere Simba is doing good, but yeah, Hamza is also doing better. So, this is the complete architecture is. If you have any question, please feel to ask. >> [music]
More from this event
See all 126 talks →
AI Is Not the Risk. Architectural Drift Is - Sunil Kalkunte
17:39
Breaking the Monolith: Tesco’s Journey to Federated GraphQL with xAPI - Vishwas Chandrashekar
29:13
A Practical Introduction to LangChain4j - Venkat Subramaniam
1:01:28
Beyond the AI Models: How Lowe’s is Building the Store That Knows - Swaroop Shivaram
13:59