{"id":12283,"date":"2023-06-29T07:00:00","date_gmt":"2023-06-29T15:00:00","guid":{"rendered":"https:\/\/formtek.com\/blog\/?p=12283"},"modified":"2022-09-06T10:37:15","modified_gmt":"2022-09-06T18:37:15","slug":"ai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models","status":"publish","type":"post","link":"https:\/\/formtek.com\/blog\/ai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models\/","title":{"rendered":"AI Chips: AI-Targeted Massive Wafers Speed the Creation of Large Language Models"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">&#8216;<a href=\"https:\/\/www.cerebras.net\/product-chip\/#:~:text=The%20Wafer%2DScale%20Engine%20(WSE,fastest%20AI%20processor%20on%20Earth.\" data-type=\"URL\" data-id=\"https:\/\/www.cerebras.net\/product-chip\/#:~:text=The%20Wafer%2DScale%20Engine%20(WSE,fastest%20AI%20processor%20on%20Earth.\">Wafer-scale engine<\/a>&#8216; (WSE) &#8211; it is what <a href=\"https:\/\/en.wikipedia.org\/wiki\/Cerebras\" data-type=\"URL\" data-id=\"https:\/\/en.wikipedia.org\/wiki\/Cerebras\">Cerebras<\/a>, a semiconductor manufacturer, calls the architecture of their massive GPU engine on a single &#8216;chip&#8217;. It is 56 times the size of the largest &#8216;standard&#8217; GPU chip and packs 3000 times more chip memory and more than 10000 times the memory bandwidth &#8212; 2.6 trillion transistors and over 850,000 cores.  The Cerebras WSE is a massive silicon chip, the largest ever built, that is about the size of an Apple iPad, and it was purpose-built to target AI.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.cerebras.net\/product-system\/\" data-type=\"URL\" data-id=\"https:\/\/www.cerebras.net\/product-system\/\">The Cerebras website<\/a> says that &#8220;A single CS-2 typically delivers the wall-clock compute performance of many tens to hundreds of graphics processing units (GPU), or more. In one system less than one rack in size, the CS-2 delivers answers in minutes or hours that would take days, weeks, or longer on large multi-rack clusters of legacy, general purpose processors.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The target of the Cerebras WSE technology is the creation of large natural language learning models, like GPT-3.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.linkedin.com\/in\/andrewdfeldman\/\" data-type=\"URL\" data-id=\"https:\/\/www.linkedin.com\/in\/andrewdfeldman\/\">Andrew Feldman<\/a>, Cerebras CEO, <a href=\"https:\/\/bdtechtalks.com\/2022\/08\/01\/cerebras-large-language-models\/\" data-type=\"URL\" data-id=\"https:\/\/bdtechtalks.com\/2022\/08\/01\/cerebras-large-language-models\/\">said that<\/a> &#8220;what\u2019s hard about training LLMs isn\u2019t the machine learning. It is the distributed compute necessitated by the GPU, the fact that it\u2019s a small compute engine. You have to break up this big problem and spread it out over lots of little engines. That work, distributed parallel computation, is obscure and it\u2019s rare and only a few organizations in the world are good at it.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/www.anl.gov\/profile\/rick-l-stevens\" data-type=\"URL\" data-id=\"https:\/\/www.anl.gov\/profile\/rick-l-stevens\">Rick Stevens<\/a>, associate director of the\u00a0 Argonne National Laboratory, <a href=\"https:\/\/www.cerebras.net\/press-release\/cerebras-systems-announces-worlds-first-brain-scale-artificial-intelligence-solution\/\" data-type=\"URL\" data-id=\"https:\/\/www.cerebras.net\/press-release\/cerebras-systems-announces-worlds-first-brain-scale-artificial-intelligence-solution\/\">said that<\/a> \u201cCerebras\u2019 inventions, which will provide a 100x increase in parameter capacity, may have the potential to transform the industry. For the first time, we will be able to explore brain-sized models, opening up vast new avenues of research and insight.\u201d<\/p>\n<div class=\"lightsocial_container\"><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/digg.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/digg.png\" alt=\"Digg This\" title=\"Digg This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.reddit.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/reddit.png\" alt=\"Reddit This\" title=\"Reddit This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.stumbleupon.com\/submit?url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F&amp;title=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/stumbleupon.png\" alt=\"Stumble Now!\" title=\"Stumble Now!\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/buzz.yahoo.com\/buzz?targetUrl=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F&amp;headline=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/yahoo_buzz.png\" alt=\"Buzz This\" title=\"Buzz This\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dzone.com\/links\/add.html?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dzone.png\" alt=\"Vote on DZone\" title=\"Vote on DZone\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.facebook.com\/sharer.php?t=&amp;u=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/facebook.png\" alt=\"Share on Facebook\" title=\"Share on Facebook\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/delicious.com\/save?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/delicious.png\" alt=\"Bookmark this on Delicious\" title=\"Bookmark this on Delicious\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.dotnetkicks.com\/kick\/?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetkicks.png\" alt=\"Kick It on DotNetKicks.com\" title=\"Kick It on DotNetKicks.com\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/dotnetshoutout.com\/Submit?title=&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/dotnetshoutout.png\" alt=\"Shout it\" title=\"Shout it\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.linkedin.com\/shareArticle?mini=true&amp;url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F&amp;title=&amp;summary=&amp;source=\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/linkedin.png\" alt=\"Share on LinkedIn\" title=\"Share on LinkedIn\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.technorati.com\/faves?add=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/technorati.png\" alt=\"Bookmark this on Technorati\" title=\"Bookmark this on Technorati\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/twitter.com\/home?status=Reading+https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/twitter.png\" alt=\"Post on Twitter\" title=\"Post on Twitter\" \/><\/a><\/div><div class=\"lightsocial_element\"><a class=\"lightsocial_a\" href=\"http:\/\/www.google.com\/buzz\/post?url=https%3A%2F%2Fformtek.com%2Fblog%2Fai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models%2F\" target=\"_blank\"><img decoding=\"async\" class=\"lightsocial_img\" src=\"https:\/\/formtek.com\/blog\/wp-content\/plugins\/light-social\/google_buzz.png\" alt=\"Google Buzz (aka. Google Reader)\" title=\"Google Buzz (aka. Google Reader)\" \/><\/a><\/div><\/div>","protected":false},"excerpt":{"rendered":"<p>&#8216;Wafer-scale engine&#8216; (WSE) &#8211; it is what Cerebras, a semiconductor manufacturer, calls the architecture of their massive GPU engine on a single &#8216;chip&#8217;. It is 56 times the size of the largest &#8216;standard&#8217; GPU chip and packs 3000 times more<span class=\"ellipsis\">&hellip;<\/span><\/p>\n<div class=\"read-more\"><a href=\"https:\/\/formtek.com\/blog\/ai-chips-ai-targeted-massive-wafers-speed-the-creation-of-large-language-models\/\">Read more &#8250;<\/a><\/div>\n<p><!-- end of .read-more --><\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[69,89,124],"tags":[],"class_list":["post-12283","post","type-post","status-publish","format-standard","hentry","category-artificial-intelligence","category-computing","category-semiconductors"],"_links":{"self":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/12283","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/comments?post=12283"}],"version-history":[{"count":1,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/12283\/revisions"}],"predecessor-version":[{"id":12284,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/posts\/12283\/revisions\/12284"}],"wp:attachment":[{"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/media?parent=12283"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/categories?post=12283"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/formtek.com\/blog\/wp-json\/wp\/v2\/tags?post=12283"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}