USER
Please, Translate the following English into Korean:
So, smooth quant makes the weight and activation into 8-bit. By smoothing the activations, pushing the quantization difficulty from the activation to the weight, improving the throughput for large language model. What about the batch size is small in the edge scenario or real-time interaction, like you want to talk to a robot, one-to-one conversation. When the batch size is small, we observe this is the real flying model. The x-axis is the compute intensity. The y-axis is the measured Tops per second. When the batch size is one, there is actually a pretty low utilization. This is your peak Tops per second. This is the measured Tops per second. It's very underutilized because it's highly memory-bounded. Because large language models are pretty big, like LAMA2-7b has 7 billion parameters. If you're using FP16, how many storage do you need? That's 14 gigabytes. That's pretty big. You have to fetch 14 gigabytes in an autoregressive manner, which means generating each token requires 14 gigabytes. So, it's a very high memory access. Remember, computation is cheap. Memory is expensive. We want to reduce those memory footprints. Compared with the weight, the activation is actually small. So the weight is 4K by 4K, but the activation, if it's a single batch, is 1 by 4K. It's a thousand times smaller. Therefore, we should focus on compressing the weight. That's why we introduced the memory. We should focus on compressing the weight. We should focus on this low-bit, weight-only quantization. We're also going to see that in the homework. So that's why we need such weight-only quantization, to reduce the memory bandwidth. Previously we had to fetch 14 gigabytes of memory. What about if we quantize the weight to 4-bit for LAMA7b? What is the memory footprint now? Three and a half. Three and a half. Three and a half gigabytes. Four times smaller. So this is naively doing that. This is a weight matrix. So using the IP16 representation, the four 8 by 4 matrix. This is the quantized version. It's again 8 by 4, but it's quantized using 3-bit, in this case, for easy to show it. But immediately we see it's perplexity degradation. The lower the perplexity, the better the quality. But here, unfortunately, there's a huge job of the perplexity, if we naively quantize it. Perplexity is a measurement of the quality of a language model. It's showing the accuracy of predicting the next word, whether you're accurately predicting the next word or not. So the lower, the better. So, unfortunately, you're actually doing to run to the nearest conspiracy to complete to a root del Mu, becauseứs andu a not exactly read each other, which is the biggest problem in computationandra mirroring the choice of theativity domain. All of the errors I gave are done byately. to nearest quantization using 3-bit hurt the accuracy a lot. Even if we're using this groupwise quantization, so here every 128 numbers are quantized together, and we have a shared scaling factor, a finer granularity of shared scaling factor every 128 elements, which we covered in the second lecture of this quantization part. Interestingly, we find not all the weights are equally important. Just by quantizing 1%, keep 1% of the rows in FP16, it helps a lot, immediately bring back the perplexity to the original value. This is so amazing, we seem to find a way to quantize that, right? Only keep 1% into FP16. Like here. Only one channel, keep it the same as before, just don't quantize it. Immediately bring back the perplexity, the quality. So therefore, we have two natural to-dos. One to-do is to think about how do we choose those salient channels, which channel is important. They exist, but the channel is only 1% of the channel. But how do we systematically, systematically select those 1% of the channels? The second to-do is, keeping FP16 will make the inference kernel difficult. How do we get rid of this mix of precision and still use full quantization? Everything will be in blue rather than having 1% in yellow. So let's answer those two questions. When we are doing pruning, how do we select these important weights? Which one to prune, which one to remove? We look at the weight itself, right? If it is large, we think, oh, that's an important weight, important channel, we should keep it. What if we do it here? We keep the weight. If it is large, we just keep it, otherwise we remove it. Unfortunately, the perplexity is pretty high. And we did it another way. Don't look at the weight. Since weight is multiplied with the activation. Let's look at those activations. Since during smooth quant we find some activations are pretty huge, we want to preserve those outlier channels. So say this is an outlier channel. This is a pretty big activation channel. And that is consistent for different inputs, different tokens. They are all big channels. And then, if this channel is big in the activation, this corresponding weight is considered salient or important. We should keep them. So using this way, activation, not the weight, that's why we call it activation-aware weight-only quantization. We are quantizing the weight. But you look at the activation to determine which weight is salient. By looking at the activation's magnitude, it's easier. To recover the perplexity. So we solve the first to-do. Look at the activation. Not the weight. To determine those 1% salient channel. Then we have the second to-do. Right? Can we not rely on this mixed precision? Still use all int 4 or all int 3 to get rid of this mixed precision. We tried a very simple technique. We just multiply this weight channel by 2 and divide the activation channel by 2. So they are mathematically equal, like smooth quant. So by multiplying it from 1.5, 2, and 4, we find this is a very effective way to recover the perplexity. Without introducing this Fp16, 1% of channels in Fp16, just multiply that salient channel by a number larger than 1 and bring back this perplexity. This is so amazing. Previously, we had to rely on 1% of the channels being unquantized in Fp16, making the kernel difficult to write. But now, just multiply the salient channel by a larger number, and it will recover. Another problem came up. How large should we multiply that channel? Here you see the perplexity decrease and the increase. There must be a sweet spot. We want to automatically search the multiplier. So let's analyze first, why enlarging the channel makes it easier to recover the perplexity? So this layer is the perplexity. So this layer is denoted by weight times deactivation. And then we care about the quantization error from the quantized variant of the W times x. So QW, the quantized variant of W, basically equal to, this is defining the range. We divide the range by the number of centroids. Since we have n bits, they are to the power of n minus 1, so we have n centroids. And this is the distance between each centroid. And then we give W divided by this number and multiply this number to the outside. And we have to round it to the nearest integer. So that's the quantized variant of W. What if we scale that? Like here, we scale it by 1.5, by 2, for those salient weights. What happens to that? We scale up the weight, and we have to scale down the activation. Previously, QW times x, now it's QW times s, times x divided by s. So they are still mathematically equal. s gets cancelled here. If we plug in Ws into the representation for the Q, we can see sw is here. Since W becomes sw. And x divided by s is here. What happens here is that the rounding function always has an expectation of 0.25. Since the rounding error ranges from 0 to 0.5. Like 1.5 gets rounded to 2, 1.75 gets rounded to 2 as well. The average is 0.25, it's a quarter. Since it's between 0 and 0.5. So this doesn't change. What about this delta? This delta is only dependent on the maximum of the weight. There are a group of weights in the vertical dimension. The group size is 128. Just scale up one channel, it's very unlikely to change the maximum value. Unless 1 in 125, 128, you hit that max value. Otherwise you are not going to change it. So this delta is not going to change. But only this s is something, s is larger than 1, so this error is scaled down. So when the s is greater than 1, the error is scaled down. That's why scaling up the salient channel can achieve the same effect of making that channel to be FP16. So one question about the scaling up. Is this scaling up after the quantization or before? Before the quantization. So it's easier to quantize. So we scaled it up. And then the equation is very similar to smooth quant. Make it easy for industry to put into products. Same infrastructure. You can do either smooth quant or AWQ. So here we times, W times s, x divided by s. This can be fused to the previous operation. Or fused into the layer norm. And then we take a data-driven approach to do a fast grid search. To search the best scaling factor, which is greater than 1. And later follow-up work even proposed a learning-based method to use gradient descent to learn the best scaling factor. So this is 3-bit group size 128 on LAMA. And also LAMA 2. AWQ shows consistent better performance compared with surround-to-nearest or GPT-Q or GPT-QR. Like 7B all the way to 30B models. It also works well for multi-model and large language model, which we're going to introduce in the next lecture on region transformers. So this is Flamingo image captioning. This is comparison with different baselines. Actually pretty significant improvement about the accuracy here. Given this image, the round-to-nearest baseline quantization model says a model airplane is flying in the sky. AWQ can say two toy airplanes sitting on a grass field. This one baseline model is saying a man is holding a gun and a baby elephant in his arm. Versus AWQ says a man and his daughter pose with an elephant. The last one, a man and a dog walking past some bushes. Versus AWQ, two dogs are working on the street. It can even use LAMA, quantized LAMA, to do visual reasoning. Like given this, there is some caption here, but this is all represented in the image format, although it has some text. You have to automate it to do OCR to understand it. Some chicken, like a world map. The baseline quantization RTM model says there are small pictures of the Earth and other planets placed on top of the food. Versus AWQ says a lighthearted and humorous take on the concept of looking at the pieces of the Earth from space. A plate of fried food, especially chicken nuggets, is presented with the caption, and the caption is actually exactly the same as the caption here. So it means this vision language model is automatically doing the OCR to understand the text here. One more example, able to recognize this, who is painting this, Leonardo da Vinci. Okay, so smooth quant and AWQ are widely used these days. That's why we also put it in the homework in the lab four, which we released last week. We give you the code, and also it will pave the way for lab five, which we are going to actually implement that on a laptop. Omidia, FASR Transformer, and TensorFlow RTLM, this is the library run large language model inference, which is actually a pretty amazing library, released actually last week. How amazing, how timely we are. TensorFlow RTLM, they are using the smooth quant and AWQ as their quantization approaches. Also Intel, who changed Berkeley's BLM, which we're going to talk about that, Berkeley Fast Chat, and Hugging Face, and SenseTime, and several open source community have been using that. Okay, so how do we translate these two concepts? So the first one is, how do we translate this theoretical saving into measured speedup? And can we deploy this large language model on edge devices, like on our laptops, our phones? Okay, so I'm going to introduce TinyChat, which is a lightweight chatbot for large language model on the edge. This is what we designed, 3D printed computer, which has an ORI Nano inside. And we also have a demo here on the right. So deploying large language model on the edge is quite useful. For example, you run CodePilot locally on your edge device, code completion, office, game chat, especially coding. Enterprise data is privacy sensitive, you don't want to upload to the cloud. But these devices are very resource constrained. Here is only a small JSON ORI Nano, resource constrained, low power, and also do not have access to the internet always. Privacy is important. Okay. So here is our TinyChat computer. It can ask you questions, and also scroll up to see the previous answers. So basically, TinyChat implements this 4-bit compressed A to the Q model. The weight is 4-bit to save the memory footprint. And here we are doing running it on different laptops. This is running the CodeLama, writing code using CodeLama. Pretty fast. And this is comparing MIT, Harvard, used to give it different prompts. Blazing fast. After Lab 4, you are going to implement something similar on your laptop. On Lab 5. And feel free to continue improving that and pushing to the code base. TinyChat engine is a very complex final project, which is open-ended. But this could be one of the choices. And a key technique to enable such fast inference is algorithm and system co-design. So on the algorithm side, 4-bit AWQ quantization. On the system side, is the Tiny Engine technique, parallel computing techniques. We introduced loop-by-routing, blocking, cache locality, right? And also multi-threading, CUDA programming, different techniques. And also how do we lay out the 4-bit weight in memory and run time decoded from 4-bit to 16-bit to avoid this decoding overhead. That is also one of the key techniques to make it run fast. This is comparing on the 1490 GPU without and with AWQ. How fast, how far, how much speed up we can get. So the FP16 version is showing both the weight and activation in FP16. So that's 50 tokens per second running on 1490 GPU. This is the AWQ version. The weight is only 4-bit. Activation is still 16-bit. Since this is the memory bottleneck, not the activation. So we keep activation in FP16 to preserve the accuracy. And we run time decode the weight from int4 by FP16 and do the arithmetic in FP16. Since computing is cheap, memory footprint is expensive. This is exactly the way we did in the efficient inference engine in ISCA 16. Now this method reborn and proved to be quite helpful to accelerate this large language model for real-time inference. This one already finished. This one is still slowly making progress. TinyChat is also flexible, support a lot of different large language models like MPT7B by Nozick, Falcon7B, and also Vicuna7B. It can also run the 13B parameter model on a MacBook. Even on a JSON-ORI, which is a mobile GPU, 30 tokens per second running on LAMA2.
ASSISTANT
자, SmoothQuant는 가중치(weight)와 활성화 함수(activation)를 8비트로 만듭니다. 활성화를 부드럽게 하여 양자화의 어려움을 활성화 함수에서 가중치로 이동시킴으로써 대형 언어 모델의 처리량을 개선합니다. 그렇다면 엣지 시나리오나 실시간 상호 작용에서 배치 크기(batch size)가 작은 경우는 어떨까요? 예를 들어 로봇과 대화하고 싶은 일대일 대화처럼요.
배치 크기가 작을 때, 우리는 이것이 실제로 비행 모델임을 관찰합니다. x축은 계산 강도이고, y축은 측정된 초당 테라연산(TOPS)입니다. 배치 크기가 1일 때 실제로 활용도가 매우 낮습니다. 이것이 최대 초당 테라연산이고, 이것이 측정된 초당 테라연산입니다. 메모리 의존성이 높기 때문에 활용도가 매우 낮습니다. 대형 언어 모델은 상당히 큽니다. 예를 들어 Llama2-7b는 70억 개의 매개변수를 가지고 있습니다. FP16을 사용한다면 얼마나 많은 저장 공간이 필요할까요? 14GB입니다. 꽤 크죠. 각 토큰을 생성할 때마다 14GB를 가져와야 합니다. 즉, 자동 회귀 방식으로 14GB를 가져와야 합니다. 그래서 메모리 접근이 매우 많습니다. 기억하세요, 계산은 저렴하고 메모리는 비쌉니다. 우리는 이러한 메모리 사용량을 줄이고 싶습니다.
가중치와 비교할 때, 활성화 함수는 실제로 작습니다. 가중치는 4K x 4K이지만, 배치 크기가 1인 경우 활성화 함수는 1 x 4K입니다. 천 배 더 작습니다. 따라서 우리는 가중치를 압축하는 데 집중해야 합니다. 그래서 메모리에 대해 이야기했습니다. 우리는 가중치를 압축하는 데 집중해야 합니다. 이러한 저비트(weight-only) 양자화에 집중해야 합니다. 이것은 과제에서도 볼 것입니다. 그래서 우리는 메모리 대역폭을 줄이기 위해 이러한 가중치 전용 양자화(weight-only quantization)가 필요합니다.
이전에 우리는 14GB의 메모리를 가져와야 했습니다. Llama7b의 가중치를 4비트로 양자화하면 어떻게 될까요? 메모리 사용량은 얼마일까요? 3.5GB입니다. 네 배 더 작습니다. 이것은 단순히 그렇게 한 것입니다. 이것은 가중치 행렬입니다. FP16 표현을 사용한 8 x 4 행렬 4개입니다. 이것은 양자화된 버전입니다. 역시 8 x 4이지만, 보기 쉽게 하기 위해 여기서는 3비트로 양자화되었습니다. 그러나 즉시 퍼플렉시티(perplexity)가 저하되는 것을 볼 수 있습니다. 퍼플렉시티가 낮을수록 품질이 좋습니다. 그러나 여기서는 단순히 양자화하면 퍼플렉시티가 크게 떨어집니다.
퍼플렉시티는 언어 모델의 품질을 측정하는 척도입니다. 다음 단어를 정확하게 예측하는지를 보여줍니다. 따라서 낮을수록 좋습니다. 불행히도 3비트로 가장 가까운 값으로 양자화하면 정확도가 많이 떨어집니다. 그룹별 양자화를 사용하더라도, 여기서는 128개의 숫자마다 양자화되고 공유 스케일링 팩터를 가지며, 양자화 부분의 두 번째 강의에서 다뤘듯이 128개의 요소마다 더 세분화된 공유 스케일링 팩터를 가집니다.
흥미롭게도 모든 가중치가 똑같이 중요한 것은 아닙니다. 전체의 1%의 행을 FP16으로 유지하면 크게 도움이 되어 즉시 퍼플렉시티를 원래 값으로 되돌립니다. 놀랍죠? 그렇게 양자화하는 방법을 찾은 것 같습니다. 단지 1%만 FP16으로 유지합니다. 여기처럼요. 단 하나의 채널만 이전과 동일하게 유지하고 양자화하지 않습니다. 즉시 퍼플렉시티, 즉 품질이 회복됩니다.
따라서 우리는 두 가지 자연스러운 할 일이 있습니다. 첫 번째는 어떤 중요한 채널을 선택할 것인가입니다. 그들은 존재하지만 전체 채널의 1%뿐입니다. 그러나 어떻게 체계적으로 그 1%의 채널을 선택할 수 있을까요? 두 번째 할 일은, FP16을 유지하면 추론 커널이 복잡해집니다. 어떻게 이 혼합 정밀도를 제거하고 여전히 완전한 양자화를 사용할 수 있을까요? 모든 것이 파란색으로 되고 1%가 노란색이 아닌 상태로요.
그래서 이 두 가지 질문에 답해 봅시다. 우리가 가지치기를 할 때, 중요한 가중치를 어떻게 선택합니까? 어떤 것을 제거하고 어떤 것을 남길까요? 우리는 가중치 자체를 봅니다. 만약 그것이 크다면 중요한 가중치라고 생각하고 유지합니다. 여기서도 그렇게 하면 어떨까요? 가중치가 크면 유지하고 그렇지 않으면 제거합니다. 하지만 퍼플렉시티는 상당히 높습니다.
우리는 다른 방법을 시도했습니다. 가중치를 보지 말고, 가중치가 활성화 함수와 곱해지기 때문에 활성화 함수를 봅시다. SmoothQuant 중에 일부 활성화 함수가 상당히 크다는 것을 발견했기 때문에, 이러한 이상치 채널을 보존하고 싶습니다. 예를 들어 이것은 이상치 채널입니다. 상당히 큰 활성화 채널이며, 다른 입력이나 토큰에서도 일관성이 있습니다. 모두 큰 채널입니다. 따라서 이 채널이 활성화 함수에서 크다면, 대응하는 가중치는 중요한 것으로 간주되어 유지해야 합니다.
이렇게 활성화를 보고 가중치를 보지 않으므로, 이를 활성화 인식 가중치 전용 양자화(Activation-aware Weight-only Quantization)라고 합니다. 우리는 가중치를 양자화하지만 어떤 가중치가 중요한지를 결정하기 위해 활성화를 봅니다. 활성화의 크기를 봄으로써 퍼플렉시티를 회복하기가 더 쉽습니다. 그래서 첫 번째 할 일을 해결했습니다. 가중치가 아니라 활성화를 봐서 그 1%의 중요한 채널을 결정합니다.
두 번째 할 일은 이 혼합 정밀도에 의존하지 않고도 모든 것을 int4나 int3로 사용하여 이 혼합 정밀도를 제거할 수 있는지입니다. 우리는 매우 간단한 기술을 시도했습니다. 중요한 가중치 채널에 2를 곱하고 활성화 채널을 2로 나눕니다. 이렇게 하면 수학적으로 SmoothQuant처럼 동일합니다. 1.5, 2, 4로 곱하여 퍼플렉시티를 회복하는 데 매우 효과적이라는 것을 발견했습니다. FP16의 1% 채널을 도입하지 않고도 중요한 채널에 1보다 큰 숫자를 곱하여 퍼플렉시티를 회복합니다. 놀랍죠? 이전에는 커널 구현이 어려운 FP16의 1% 채널에 의존해야 했지만, 이제는 중요한 채널에 더 큰 숫자를 곱하기만 하면 됩니다.
하지만 또 다른 문제가 생깁니다. 그 채널에 얼마나 크게 곱해야 할까요? 여기에서 퍼플렉시티가 감소했다가 증가하는 것을 볼 수 있습니다. 최적의 지점이 있어야 합니다. 우리는 곱셈 인자를 자동으로 탐색하고 싶습니다.
먼저 분석해 봅시다. 채널을 확대하면 왜 퍼플렉시티를 회복하기 쉬워질까요? 이 레이어는 가중치 곱하기 활성화 함수로 표시됩니다. 우리는 양자화된 W 곱하기 x에서의 양자화 오류에 관심이 있습니다. 양자화된 W(QW)는 W의 양자화 버전입니다. 우리는 범위를 정의하고, 센트로이드 수로 나눕니다. n 비트가 있으므로 2의 n승 마이너스 1개의 센트로이드가 있습니다. 각 센트로이드 사이의 거리는 이렇습니다. W를 이 값으로 나누고 바깥쪽에 이 값을 곱합니다. 그리고 가장 가까운 정수로 반올림합니다.
만약 우리가 이것을 스케일링하면 어떻게 될까요? 중요한 가중치에 대해 1.5나 2를 곱하면요. 우리는 가중치를 스케일 업하고 활성화 함수를 스케일 다운해야 합니다. 이전에는 QW 곱하기 x였지만 이제는 QW 곱하기 s, 곱하기 x 나누기 s입니다. 수학적으로 동일하며 s는 상쇄됩니다.
우리가 W에 s를 곱한 것을 대입하면 sw가 나오고, x를 s로 나눈 것도 나타납니다. 여기서 반올림 오류는 항상 0에서 0.5 사이이므로 평균은 0.25입니다. 따라서 이것은 변하지 않습니다. 델타는 가중치의 최대값에만 의존하며, 하나의 채널만 스케일 업하면 최대값이 변할 가능성은 매우 낮습니다. 따라서 델타는 변하지 않고, s가 1보다 크므로 오류는 줄어듭니다. 그래서 중요한 채널을 확대함으로써 해당 채널을 FP16으로 만드는 것과 동일한 효과를 얻을 수 있습니다.
스케일링은 양자화 전에 이루어집니다. 이렇게 하면 양자화하기가 더 쉽습니다. 그리고 이 방정식은 SmoothQuant와 매우 유사하여 산업에서 제품에 도입하기 쉽습니다. 같은 인프라를 사용할 수 있습니다. SmoothQuant나 AWQ를 사용할 수 있습니다. 여기서 우리는 W에 s를 곱하고 x를 s로 나눕니다. 이는 이전 작업에 통합되거나 레이어 노름(layer norm)에 통합될 수 있습니다.
그리고 우리는 데이터 기반 접근법을 사용하여 1보다 큰 최상의 스케일링 팩터를 빠르게 그리드 검색을 통해 찾습니다. 후속 연구에서는 그레이디언트 디센트를 사용하여 최상의 스케일링 팩터를 학습하는 방법도 제안되었습니다.
이것은 Llama와 Llama 2의 3비트 그룹 크기 128에 대한 결과입니다. AWQ는 가장 가까운 값으로 반올림하거나 GPT-Q 또는 GPT-QR과 비교하여 일관되게 더 나은 성능을 보여줍니다. 70억에서 300억 매개변수 모델까지요.
또한 다음 강의에서 소개할 멀티모달 대형 언어 모델에서도 잘 작동합니다. 이것은 Flamingo 이미지 캡션입니다. 다양한 베이스라인과 비교하여 정확도가 상당히 향상되었습니다. 이 이미지를 주면, 가장 가까운 값으로 양자화한 베이스라인 모델은 "모형 비행기가 하늘을 날고 있다"고 말하지만, AWQ는 "잔디밭에 놓인 두 대의 장난감 비행기"라고 말합니다.
또한 Llama를 양자화하여 시각적 추론도 할 수 있습니다. 예를 들어, 이미지를 주고 자동으로 OCR을 수행하여 텍스트를 이해합니다. 베이스라인 양자화 모델은 "음식 위에 지구와 다른 행성의 작은 그림들이 놓여 있다"고 말하지만, AWQ는 "우주에서 지구 조각을 보는 개념에 대한 가벼운 유머러스한 접근. 특히 치킨 너겟으로 구성된 튀긴 음식 접시에 캡션이 제시됩니다"라고 말합니다. 캡션은 실제로 이미지에 있는 캡션과 정확히 동일합니다. 즉, 이 비전 언어 모델이 자동으로 OCR을 수행하여 텍스트를 이해한다는 것입니다.
또한 누가 그린 그림인지 인식할 수 있습니다. 레오나르도 다빈치죠.
요즘 SmoothQuant와 AWQ는 널리 사용되고 있습니다. 그래서 우리는 지난주에 출시한 과제 4에 그것을 포함했습니다. 코드를 제공하고, 과제 5에서는 실제로 노트북에서 그것을 구현할 것입니다.
NVIDIA의 FasterTransformer와 TensorRT-LLM은 대형 언어 모델 추론을 위한 라이브러리로, 실제로 지난주에 출시된 놀라운 라이브러리입니다. 이들은 SmoothQuant와 AWQ를 양자화 접근법으로 사용합니다. 또한 Intel, Berkeley의 BLM, FastChat, Hugging Face, SenseTime 등 여러 오픈소스 커뮤니티에서도 사용되고 있습니다.
자, 이 두 가지 개념을 어떻게 구현할까요? 첫 번째는 이론적인 절약을 실제 속도 향상으로 어떻게 전환할 수 있느냐입니다. 그리고 이 대형 언어 모델을 엣지 디바이스, 즉 노트북이나 휴대폰에 배포할 수 있을까요?
TinyChat을 소개하겠습니다. 엣지에서 대형 언어 모델을 위한 가벼운 챗봇입니다. 우리가 설계한 3D 프린팅 컴퓨터로, Jetson Orin Nano가 내장되어 있습니다. 오른쪽에 데모가 있습니다.
엣지에서 대형 언어 모델을 배포하는 것은 매우 유용합니다. 예를 들어, 코드 자동 완성(CodePilot)을 로컬 엣지 디바이스에서 실행하여 코드 작성, 오피스 작업, 게임 채팅 등을 할 수 있습니다. 특히 코딩 분야에서요. 기업 데이터는 프라이버시가 중요하므로 클라우드에 업로드하고 싶지 않습니다. 그러나 이러한 장치는 자원이 매우 제한적입니다. 여기서는 작은 Jetson Orin Nano로, 자원이 제한적이고 저전력이며 항상 인터넷에 연결되지 않을 수 있습니다. 프라이버시는 중요합니다.
여기 우리 TinyChat 컴퓨터가 있습니다. 질문을 할 수 있고, 이전 답변을 스크롤하여 볼 수 있습니다. 기본적으로 TinyChat은 4비트로 압축된 AWQ 모델을 구현합니다. 메모리 사용량을 줄이기 위해 가중치는 4비트입니다.
여기서는 다양한 노트북에서 실행하고 있습니다. 이것은 CodeLlama를 실행하여 코드를 작성하는 것입니다. 꽤 빠릅니다. MIT와 하버드에 대한 다양한 프롬프트를 비교합니다. 매우 빠릅니다. 과제 4를 마치면, 과제 5에서 노트북에서 이와 유사한 것을 구현할 것입니다. 계속 개선하여 코드 베이스에 기여하셔도 좋습니다. TinyChat 엔진은 매우 복잡한 최종 프로젝트로, 개방적입니다. 이것이 선택지 중 하나가 될 수 있습니다.
이렇게 빠른 추론을 가능하게 하는 핵심 기술은 알고리즘과 시스템의 공동 설계입니다. 알고리즘 측면에서는 4비트 AWQ 양자화입니다. 시스템 측면에서는 TinyEngine 기술, 병렬 컴퓨팅 기술입니다. 루프 펼치기, 블로킹, 캐시 지역성, 멀티스레딩, CUDA 프로그래밍 등의 기술을 도입했습니다. 또한 4비트 가중치를 메모리에 어떻게 배치하고 런타임에 4비트에서 16비트로 디코드하여 디코딩 오버헤드를 피하는지도 중요합니다. 이것이 빠르게 실행하는 데 중요한 기술 중 하나입니다.
여기서는 1490 GPU에서 AWQ를 사용하지 않은 것과 사용한 것을 비교하고 있습니다. 얼마나 빠른지, 얼마나 큰 속도 향상을 얻을 수 있는지 보여줍니다. FP16 버전은 가중치와 활성화 함수가 모두 FP16입니다. 따라서 1490 GPU에서 초당 50 토큰을 생성합니다. AWQ 버전에서는 가중치는 4비트이고 활성화 함수는 여전히 16비트입니다. 이는 메모리 병목 현상이 활성화 함수가 아니라 가중치 때문입니다. 정확도를 유지하기 위해 활성화 함수를 FP16으로 유지합니다. 런타임에 가중치를 int4에서 FP16으로 디코드하고 FP16에서 연산을 수행합니다. 계산은 저렴하고 메모리 사용량이 비싸기 때문입니다. 이는 우리가 ISCA 16의 효율적인 추론 엔진에서 했던 방식과 정확히 동일합니다. 이 방법은 부활하여 실시간 추론을 위한 대형 언어 모델 가속에 매우 유용하다는 것이 입증되었습니다.
이것은 이미 완료되었고, 이것은 아직 천천히 진행 중입니다. TinyChat은 또한 유연하여 MPT7B, Falcon7B, Vicuna7B와 같은 다양한 대형 언어 모델을 지원합니다. MacBook에서도 130억 매개변수 모델을 실행할 수 있습니다. 심지어 모바일 GPU인 Jetson Orin Nano에서도 Llama2를 실행하여 초당 30 토큰을 생성할 수 있습니다.