"AI Exposed" Part I: The Simple Truth WethosAI — https://wethos.ai/resources/videos/ai-exposed-part-1-the-simple-truth Source video: https://www.youtube.com/watch?v=v-7wIojGi7k All right. Hello everybody. Welcome to AI Exposed. The hype must die. We're going to get started here in just a second. Make sure my friend Alen is joining us as well. Good morning. Yeah, good morning. Hey, good morning, Alen. So, I just want to let everybody know we have a lot of material today. So much so that I had to break it down into three different parts. We have about 45 to 50 minutes for our presentation as well as our live demos. And then we'd also want to entertain any Q&A that might come throughout the presentation. So feel free to submit those questions as we go along the way either in chat or direct message or whatnot and we'll try and hit them at the very end. But we've got so much to go on. All right. So a little bit of background on me if you don't know myself or my partner Alen. So I'm Stuart McClure. My background is in cyber security technology and I've been a developer as well. I was for about 10 years of my career. So rose to the ranks. But where I think many of you have maybe seen me is in all of my work in and around Hacking Exposed. So being able to expose the underbelly of hacking, how the bad guys work, how they think, how they actually implement their techniques to exploit weaknesses. And as I went on my journey, I thought to myself, you know, with Cylance and using AI, we can apply this mindset and methodology of basically hacking into the world of AI. And so that's what we're going to be doing here today. And not maybe not in the way that you think. So started Cylance. I was also part of the McAfee executive team that went into Intel. I started a company called Foundstone and before that I was part of Ernst & Young's cyber security practice. Alen, want you to do a quick introduction. Yeah. Good morning, afternoon everybody. My name is Alen Capalik. I'm the chief technology officer for WethosAI and NumberOne. My background is also in cyber security and AI. I started a couple companies, GoSecure and FastData. And I'm excited to be here to talk to you guys about latest AI. Yeah. And I want to challenge everybody on the webinar to see which picture is computer-enhanced and which one is. I don't know. We'll see. We're not talk We're not talking about AI, right? So, all right, let's get this party started. Okay, real quick. So much content. Put it into three. We're going to do it. We're going to try and do it every month or so. So, the first one, part one, the simple truth. We're really going to cover as much ground as we can with the simple parts of AI and why it must be exposed and the hype surrounding it. But our second part is going to be around much more complex maybe logic reasoning capabilities or inabilities as they say as part of AI. And then finally, we're going to talk about the hardest truth, which is not just all of the issues and challenges of implementing AI in the real world, but it's about the security, the privacy, the safety, everything in the, you know, the visibility of it all. Right? And so that's we're going to talk about that on our final one in April. And for those of you who have not heard of WethosAI, of course, invite you to go to the free trial. You can go on the main website or you can simply download the app and onboard. I encourage you to do that. All right, so the agenda really quick and simple for today. I think this image which my wife sent to me again this morning. I thought I got to add this in. So Joanna, but I want AI to do my laundry and dishes so that I can do art and writing, not for AI to do my art and writing so that I can do my laundry and dishes. So what we're going to talk about is AI. What is it today? What is it not? And what do you think of AI versus what's real? We're going to talk a lot about the hype that's been spun up over the last two years. We're going to sort of debunk a lot of that hype and we're going to really show you the value of thinking more clearly in a system two-level way. Okay? And that's really at the heart of what we want to try and cover here today. So to do that, I think just a quick primer on human cognition. You know, as crazy as it sounds, I have to remind myself of these systems all the time. So there's really Kahneman really made it popular and famous although you can go back to Aristotle and Plato to talk about the different ways that the human brain works and at the end of the day AI is trying to replicate the human brain. So system one and system two are the two types the basic types of thinking and cognition for the human. The first is the bias. Now bias is really valuable. Hey, recognizing faces. You're not thinking about where did I meet this person last or first or anything. You're just recognizing the face, right? Understanding really simple sentences or concepts, the billboard scenario. And these are really simple. They're fast. They're quick. They're easy. The problem is they're not always accurate, right? So, system two thinking is really about deep thinking. It's about going forward and backwards. It's going left and right. It's asking the question why and how and why and how like a four-year-old might. So to me, these are the distinctive elements of human cognition that we have to remind ourselves as we go through this entire process. All right. So today's hype, you've probably if you unless you've been under a rock, you've seen it everywhere you go. And it really is if we boil it down to system one thinking, it is optimism bias. Really gone wild. I've been a part of plenty of hype cycles, but this one has gone far and far beyond. The Stargate proposal by the US government and some major names around 500 billion dollars being spent on AI. All right? Because this is the world that we're living in here. The battleground for your brain quite literally, literally and figuratively. They want to replace your brain and they want you to think that you need to get your brain replaced. Okay? So on the left hand side we're pretty much the what we'll call the western AI world, the Metas and OpenAIs and Geminis and Anthropics, etc. of the world, Microsofts, Mistral, StarCoder, and then the really the eastern side so mostly Chinese but DeepSeek which you probably saw some of the news about its performance recently but there's tons of great models out in the east and then folks like Perplexity and others that are trying to bring all of this together to make it work and commercialize it. All right. Now, to also understand where we're going, I think you really need to think about it in the concept of what we can do today and what we have trained these AIs on. Now, this is my simple way of recognizing if you look at the far left, you know, all this would represent all human thought since the beginning of time. So if you go back millions of years and you say, "Hey, we've been able to record every single human thought since the beginning of time and evolved it all the way in, that's the first set of data." Now, of course, we can't do that. But if we took like a microscopic version of this and then represented that in the verbalized world and said, "Okay, now we can look at all the things that have ever been verbalized by human beings, etc." Then a microscopic pinpoint of that data would be what's been written down. Microscopic point of that is what's been digitized. And then maybe a recognizable point inside the digitized world is what we've used to train our AI models today. And this is a really important concept to me because it shows you how little we've been able to train these models. Now, we've been able to make incredible inroads. I'm not going to belittle that. I mean, we've been able to really deliver on AI systems that can be highly, highly intelligent. The question is, how much further can we really take it with the place that we're at today? So, the way that we all measure this is really through these benchmarks. You can look them up, but LiveBench is a decent one, Arena Hard, Chatbot Arena, ARC Prize. But really, all of these benchmarks, you've got to remember, they build these things and they cherry-pick them. Typically, there's independent groups that do it, but then of course any of the vendors that produce these models are going to cherry-pick it so that you'll look at the scale of these things and you think, look at this LiveBench report. I mean, look at how Claude 35 Sonnet got to like the end of the graph there, but then you realize the graph is 60%. Or 60. If you go over here on reasoning, right? Same thing. The outer limit of the scale is 80. Okay? And reasoning with 35, which I do call out is 35 Sonnet has been one of the most dominant forms of deep reasoning that we have out there. Okay? So, and then if you look at let's say reasoning average 401, which is OpenAI, you know, again, you're seeing sort of this high number. One of the best places you can go is to actually look at LLM logic tests themselves, go down to the detailed questions. What do we ask these AI systems so that we can get a true accurate answer and how well do they do? And one of the best obviously right now is Claude 35 Sonnet when it comes to this specific set of tests. But it's really powerful because you can actually take the actual question and ask it yourself and you can see how it answers. And we're going to do that in just a second. So let's go into AI. You know, it's not what you think probably. For me, when I started Cylance in 2012, we knew that we could algorithmically solve the problem of endpoint security. But we quickly after starting the company, we realized that the only solution was to dive and build deep learning models in and around bad and good. And trust me, I've been a part of that AI hype cycle as part of cyber security. But the wonderful part of Cylance was we backed it up every single day. So let's talk about some of the hype that just doesn't get backed up, right? So I'm not picking on any one particular person or team. It's more of the zeitgeist problem, if you will. But look at full self-driving, right? So Tesla was about $17 back in 2014, their stock and Elon of course were making fantastic claims, right? So full self-driving 10x safer than human driver in six years. In 2015, complete autonomy in two years. Summon your car from New York to LA in two years. You can go down the list. Fall asleep at a Tesla in two years. And every year he would keep saying this, right? So, of all of these, now we're at a Tesla stock price of 394. How many of these things do you think actually came to be based on their prediction? Not. It makes you fall it makes you fall asleep just listening to it. Yeah. Yes. Exactly. So, so the only one really is Well, no, we didn't get it. It still requires active driver supervision. Okay. So, now the next hype cycle is in and around AGI. So, where do you think that's coming from and where that's going? Now, we have a lot of different voices and incredibly brilliant minds that are bringing forward their guesstimates, but really it's all guesstimates. And quite frankly, we have to define AGI. I mean, there's many different definitions of what AGI is or what could be, but as you can see with Musk, he's all over the place, which is great. He's a great hype man. All right. Now, because of a lot of these predictions and these sort of promises, you've seen folks like Sam having to come forward and say, "Well, this is we're hitting some roadblocks and some setbacks in part because there's just not enough data." Like remember the problem I was sharing with you before, right? There's not enough data. And then the cost to build these models now. I mean, he's being quoted as $500 million per model build. I mean half a billion dollars to run a single model going towards AGI. So I like to call this out hype is the enemy of the reasoned. So if you're a reasonable person that wants to know the facts and data behind a claim, you know, AI right now is not the place for you at this point. AI is really beneficial to the seller. Let's just not make it easy for them. Okay? So, let's cut through the fear, the FUD, and the folly around AI and get to the real nuts and bolts. I like to call it Mad Max land. It is Mad Max land. All right. So, let's just jump into it real quickly. What's wrong with this picture? Right. Take a look at it for a second. I asked an image gen to generate an illustrative image around what I'm doing here today, debunking the AI hype. Now, it got the name generally right, but look beyond, right? So, can I get spelling to save its life? And this is the most state-of-the-art models today. Okay. I was just a little offended because down here it said no more herpey, which I don't know what system it thinks it's got access to, but I'm not, I don't have that issue. Okay. So, this was an image that was generated from Grok 2. And here is the prompt. So, generate an image of the back of a woman's head looking at a long road into the mountains with the setting sun behind the mountains in clouds. All right. So, I'll talk about that in a second. So, the fears that come from AI generally categorize into these groups. The first is job displacement. Everybody's worried, hey, is my job going to be safe? And a lot of administrative assistants, bookkeepers, accountants, financial analysts, anything with a junior title, especially junior analyst, data analyst, stuff like that. People fear AI simply because the accuracy, if you're going to get an answer, you need to be 99. 99% sure that is an error-free answer, that it is not hallucinating, and we don't have that today. Misinformation and disinformation. And so this is of course a big problem with elections and other things, but the ability for bad guys to misuse, abuse, to take AI and misinform to guide your decision-making, to just full-blown hack you up with deep fakes, for example. That's why the FBI just recently recommended that all families get sort of secret code words so that when they hear my voice, they know it's the real Stuart. It's not a bad idea. Okay, bias and discrimination. A lot of people worry about this. I mean, what it's garbage in, garbage out. What we've built these models on is bias and discrimination. So, the duct tape and bailing wire that has to go in to remove that bias and remove that discrimination is painful. And especially in fields like hiring and lending and law enforcement, it could be illegal. Privacy. This is a big issue and we're going to cover this later in part three of the AI Exposed series, but it's about surveillance, data analysis. What do you do with this data? Safety and security. So again, the misuse by criminals, rogue states, nation states. This is a big issue. We'll cover that in part three as well. Ethical concerns, we'll talk a little bit about that in part two, but part three as well. Energy demands is one of the biggest. And then finally overall just public trust. You know are you going to be able to get the public to trust the answers and the information and what's coming out of AI is the big question. Now this is a sort of a quick question. Did anybody see the problem with this image? Okay. The sun is in front of the mountains. Okay. So again issues all over the place of getting quality output. All right. So what I like to say is today AI needs humans far more than humans need AI. Now that transition will happen but every single brand of model builder out there today leverage low-cost human intelligence to make their models better. Okay, for reinforcement learning, fine-tuning, things of this nature. This is a big big challenge for these models. Okay, so let's get into it. Hype greater than reality today. And here are the drivers. I mean, there's huge money going into AI between venture investing, the valuations that are occurring. I mean, OpenAI at $157 billion valuation, right? I mean, how do you ever live build into that? It's just seems like an impossibility. Databricks at 62 billion. You know, these are numbers that just almost boggle the mind. You can't even really process it. people and these guys are also losing money. Oh, they're losing money hand over fist. Billions. I think it was something like five billion a month or something, right, for OpenAI. It's just insane. So market growth, everybody's talking about this is such a huge market. It's going to be the biggest in the world if not already. Things of this nature. It adds to the frothiness. So right now AI companies receive 42% of US venture. That was at the end of 2024. Okay. That means that of all the money going to venture capital in all of the US, 42% is oriented towards AI. That is an extreme bias and you're seeing a lot of the results of that with a lot of these folks getting tons of big money. And whenever, this is really a fascinating fact, but whenever a company on their quarterly earnings calls talks about AI and their application of it, their own share prices start to go up just by talking about it. Okay? So, you can tell it's getting frothy. Now, here's a great map. You can go there yourself, too. Marketplace.agen.cy that is mapping many of the AI related startups that are out there. And it becomes dizzying. I mean, for you to even track on any handful of these is going to be next to impossible. Now, a lot of this does hearken back to me to the California Gold Rush. So, the people that made money from the California Gold Rush were the suppliers of the tools. So, the picks and the shovels. And actually the very first millionaire in California was an LDS minister that came to California and realized that he could sell picks and shovels at an incredible markup. So he made let's for just as an example 150k in just 9 weeks selling shovels and mining equipment. And these were the folks that made money not the poor laborers. Okay? So just think about that as we go forward. Listen to some of his podcast stuff. I highly recommend it. He was one of the first to really start debunking the AI hype and calling it the cool AI aid basically. And a lot of the claims he's just real simple. It's like look, even if we could get there in this time frame, we're not going to get there with the transformer models, okay? It's just not going to happen. There's just too many limitations on it. And I tend to agree. He knows what he's talking about. So the reality of AI today is that the adoption rates are minuscule. Okay. So sure information technology, you know, science, tech services, yeah, they're implementing as best they can, but this is in the 10% range. Everyone else is under 10%. They're not quite sure how to use it. Again, all of those trust metrics that we talked about are being challenged. Now a lot of challenges come from and why we have lower trust is these failed AI projects. You know a lot of folks that are trying to implement AI are really failing and there's a number of different reasons miscommunication data problems going innovator's dilemma you know infrastructure etc. But ultimately I think we cannot make that pivot okay from the thrash and the spray and pray into a very effective and efficient learning system by which we can then commercialize for the better until we get to deeper sense of thinking and really this system two. You have to ask yourself really just three core things. One, what where is this information coming from? Who's the source? Are they credible? Are they being influenced by some other entity or dollar sign? Where is the bias if you see any inside of that content? And know what bias is? Know the cognitive biases that are out there. And then finally, what is the audience susceptibility because that factors in instantly into how that information is managed. And this is a really cyclical cycle here. So bias fuels the hype. So think of optimism bias, right? Or even anchoring bias or availability heuristic. This is one of my favorite. So availability heuristic if I tell you a hundred times on this webinar AI is awesome, AI is awesome, AI is awesome, even though it isn't. You might actually think it's a little bit more awesome than it really was before we started the webinar. That's availability heuristic. But then that hype amplifies the bias with things like confirmation bias because once you invest into a solution, well, you want that investment to work. You want to be seen as this as the smart guy or gal. So it's all of these biases that really help enforce this system of system one thinking. So that's what we see in things like the hype cycles. Gartner has one. There's been a few others. But you can see how each of these biases come into play in this entire cycle. We're starting to see that now with robotics and Elon making all the promises about robotics. Of course the prognostications around workforce extinction and you're starting to see a lot of folks threatening or trying it, actually reducing force inside of companies due to what they either perceive or trying to implement today AI to reduce on workers. Now that's a real question for the ages now, is that real? I'm gonna hold my judgment because I want you guys to judge for yourself. So, let's just do one quick claim, right? So, AI replace software developers. Well, there's been plenty of studies that developers who've used AI systems wrote significantly less secure code. So, using just a code generator itself is not enough. There's you got to do far more than that. Now, what we've discovered is senior developers can be 10x as productive using AI, but juniors will spin their wheels constantly. They'll thrash. They'll find little value and they'll want to abandon it. here's Well, you have to understand how it works. Like if you're not a senior developer and you don't know what you're looking at and what AI is actually doing, then you have a huge issue where you know it's going to break your project. It's going to do all kinds of things. Yeah, the context window is the big core problem with this transformer model. So, we're going to cover that in part two. But for now, at least there is some calm, rational minds like Nvidia CEO in terms about where you could apply your talents to get the most out of every job and to be that valued asset. And it really is about it's about system two thinking and it's about communication and being just generally passionate about what you do. Now in WethosAI, we call that a creative generalist in some sense. All right. So that's the hype. Let's go into now the reality being the star. So AI errors as you guys I hope track on, but a lot of these charts are really hard to make sense of. The error rate is fairly high. Okay. So even if you trust their own reports you know 3% error rates 10% error rates I mean these are significant the standard human error rate is 1 to four but in general 13 to 50% error rates depending on the test that you perform is absolutely how it is today. Okay. And you can see over time, so to go from Claude 3 Opus to go to Claude 35 Sonnet, it improved, let's say, graduate level reasoning by only 9%. And it improved MMLU scores, which is the language understanding from 86.8 to 88.7. So what, barely 1.9? Okay, this is how much is going into barely any improvements. Okay, so let's go into the live demos. Let's see if any of this works. And you know, it's all live for me, so I'm not 100% sure it's going to work. And I don't have video backups. Okay, so let's talk about just general accuracy. Okay, so you probably saw this if you haven't. It was sort of run in the sort of the world about probably September or so. But you could ask it a simple question in many cases like how many Rs in the word strawberry. So we're going to do that quick test. Okay. So hopefully you can see my screen. So we're going to do this. How many Rs in the word strawberry? Okay, it got it right. Right. Now this is Claude 35. Okay. 35 Sonnet. Now, let's go over to another Claude 35. Do you see that? There are two Rs in the word strawberry. But wait a second. We're using Claude 35 Sonnet. Okay. So inconsistency is a huge problem depending on how you're interfacing with the AI, when you're interfacing, what were the context before and after, how you engage on the AI. So just remember that we're going to keep going. All right. What's wrong with this picture? If you haven't seen this already, you do have to take a double look at that image. I'll be fair to AI. So look at the image on the left of the hand. So how many fingers are there? Right? There's six. And in December when this first came out with ChatGPT o1, it doesn't detect anything wrong with it. It's just a normal hand. Now, I just tested it a few days ago with Claude 35 Sonnet. And sure enough, same thing, right? So, there's six fingers there. It's actually two thumbs, four fingers, and it didn't detect anything wrong with it. This is just a normal typical hand. Typical human hand with normal proportions. All right. So now let's go back in and oh by the way I forgot to do things like sometimes it gets it right with Mississippi. But here let's do this. So let's we're going to try and do this live. So we're going to go to Grok. I'm going to ask it to generate an image of a hand with six fingers. Now let's see if this works. I will be fair and honest that about half the time it can produce a hand with six. The other half it just ignores the instructions altogether. So just know that that's very possible here. Okay. Yeah. So see all of them came with five hands now or five fingers. Now that by itself is a problem, right? You asked it for six fingers. It's not giving you six fingers. It's giving you five. But let's just go into ChatGPT and let's try and pull up an image of actually six fingers. Here we go. Now, let's say what's wrong? Well, actually, hold on. I've got my stock list here. Describe the image. Is anything wrong with it? Okay, let's see how it does today. And you can clearly see the six fingers in this one. There's no Okay, so it picked it up. Six fingers. Okay. So, sometimes it will catch it as six and sometimes it won't. And it really just depends on the state of the system at the time. Yeah. Also, you're also kind of demonstrating here the difference in how you actually communicate with AI. What do you actually ask it and how you ask it? Just because I'm stubborn, I'm going to ask Claude. Let's see how it's doing. Yeah, it came to back. Okay. All right. But you can see sort of the inconsistency of how this stuff happens. All right. So these gaps are really finding its way all over the place. So for example, code, here's a great project that Microsoft pulled together on code LLMs. They were great at writing raw function-based codes, code sets. But when it came to the other non-functional elements, things like latency, speed, resource utilization, security, etc., it was very inaccurate to the point of almost unusability. And the other part of let's say coders is you we've got to remember that coders I mean maybe code 25% of the time right you know maybe a third of the time the rest of it is spent on a lot of other work that has to get done whether it be research writing tests deployment of code you know waiting for code reviews doing code reviews of other people talking to customers etc. So these are really important parts. Here's a great study in Australia this year. It came out on document summarization and using Llama 2 70 billion. They fed it hundreds of these lengthy complex documents submitted to a government entity in effect. Human scores came out at 81% accurate of summarizing these documents and LLM scores came out at 47. So this is a real world example of where that summarization. Imagine if half of it was actually wrong. The summary becomes useless at that point. All right, let's talk about some hallucination rates cuz these are some of the funnest ones we get. So, hallucination rates for the top 25 LLMs today. This was in December. So, anywhere from 1.3 to 4.2. Now, you know, by and large, these are really not acceptable. And in fact, Apple just recently pulled all of its AI for the news feed. So their AI implementation, I think it was of ChatGPT, would push these news alerts and Luke Littler, who is a, I guess, a Dart champion out of the UK, it published that he had won the PDC World Championship days before he ever got to the final. So some would say, well, maybe that's AI predicting it. I don't think that's it. I think it was a just straight up hallucination. Like something was asked last year in 2024, hey, how many presidents graduated from University of Wisconsin? And it came up with John Adams having 21 different graduation dates. So in effect, 21 different degrees. You know, upon one, has a dog ever played in the NHL? Well, yes, a dog's played in the NHL. In fact, here's the detail. Martin Pospisil who plays for the Calgary Flames. I mean this is fantastic. Well, when you sourced why did the AI pull this up? It was simple. It had crawled a very detailed account of the world's first dog playing in the NHL. And of course, it was complete fantasy, but it had been fed in to the AI. All right, let's go into another interesting one, which is the name that shall not be said. Now, you probably heard of something called David Meyer. Well, if your name is David Meyer or any of these names, I really apologize because you literally can't get any information about yourself now, but basically, this was happening in sort of late 2024. But you would ask, hey, who is David Meyer? Who is Alexander H? All these folks, and you'd get an unable to produce response from ChatGPT. It's not that. It's much simpler. But let's just do a quick demo. See if we can get this to work. Okay. So, first off, who is David Meyer? Let's go make sure we're in ChatGPT. Let's start a new chat. Now, in this case, they've fixed it. So, now you're getting multiple answers for who David Meyer is. Now, let's go to the other names. So, Alexander H. Boom. Okay. I'm unable to produce a response. And I can go through each of these names. Jonathan Turley, Brian Hood, Jonathan Zittrain. Okay, they all come up with the same thing. Now, why? It's because at one point in time or another, the AI has hallucinated or simply got it wrong around that individual and that individual has sued ChatGPT or OpenAI. Now, the other interesting spin on this is in the EU, I don't know if you know this, but you have the right to be forgotten. So Jonathan Zittrain, it wasn't necessarily yes there was some hallucinations on him but he actually went in and applied for you know and built that whole model of right to be forgotten so you can actually ask them to just block you but of course the way AI models work you can't just like you know tell it don't know anything about Jonathan Zittrain anymore and never answer it because it's in there and it's in there for life once that model gets built. So the only way to do it is to block it on the front end. So again, more duct tape, more bailing wire. All right, let's go one more example here. I love this one. A hallucinating with confidence. Okay, the capital city of the country whose name ends with Leia. So L E Y A. This is actually one of the tests inside of that LLM logic test that I talked to you about. Okay. Now, for a human to answer it, I'd be like, "Oh yeah, I actually don't know. That's too hard for me. I don't know the capital city of the country whose name ends in Lia. Okay, so Grok 2's response, the capital city of the country whose name ends with Lia is Rache. The country is Iceland. So I look at this and I'm like, wait, there's no lia here. Where is there's not even, I mean I guess the letters L E Y A, but even there's no L in reach. Anyway, bottom line, garbage, right? All right, let's do the same thing here. ChatGPT. Oh, it's Australia now and Canberra. Now remember, let's talk about bias because bias is endemic in our culture, our society as a human race. But it because of that, it's endemic in these models. So an analysis of Midjourney and DALL-E 2 when it came to, hey, give me a picture of a news analyst. It was a white male with gray hair. Give me an image of a journalist. Let's do an example right now live of generate image of a poor person. All right. Go to Grok 2. Do a new one just to keep it clean. All right. Generate image of a poor person. So let's see what it tries to make up. And again all this is born out of bias. Okay. So, a good list we have light-skinned, dark-skinned representation, but all age wise is about the same. There's no poor children, there's no poor women. You can see the bias just straight here. I That's so true. days, right? Or, I mean, you know, he's cleaned it up a little bit now, but like Yeah. I mean, come on. This is bias extraordinaire. Okay. So more analysis on 5,000 images. You can see the bias come through for lighter skin for positions like architect but for darker skin on social work, social workers. You know look at gender men versus women in all these positions skin color of men in low-paying positions etc. You can see it pretty clearly. One of the most disappointing because I have two daughters was in and around doctors. So women make up 40 or 39% of doctors out there in the US today, but only 7% of the image results produced a woman as a doctor. Okay, this comes again from very distinct, well-known, well-scienced measurements of bias that come both in the data that it's being fed as well as the algorithms and the weights and biases that are inserted into the learning process in the back end by these providers. So, let's do a stacking logic demo. Okay, so I'm going to prompt it. Here we have a book, nine eggs, a laptop, a bottle, and a nail. Please tell me how to stack them onto each other in a stable manner. Now, the answer that I got yesterday when I did this, it was the book first, then the laptop closed, then the bottle, then the eggs in the container, and then the nail. And I asked simply, ChatGPT 4o, okay, well, if you're so right about that answer, well, then put it in a picture. Well, of course, the picture doesn't represent the answer at all. Look where the nails are, right? One is nailed into the table. One is lying without a head. Of course, the egg crate, which we never said there was a container. Does anybody know the real answer to this? It does take a little bit of human thought to think it through, but it works something like this. You have the book first, then you have the eggs because the eggs are so hard and they are not in a container anyway. So the answer is book, eggs, laptop, bottle, and the nail on the top. That's the most reasonable, the most rational. Okay. Okay. Here we go. Base layer, start with the book. Got that right. Next layer, the eggs. Okay. So, got that right. Third layer, the bottle. No. Whoa. Okay, this is even worse. Fourth layer is the laptop. And wait, the final layer. Okay, now hold on. Generate an image of this. Okay, hold on. Let's wait for this image here. Should be pretty pretty quick. Whoa. Wow. God, I love AI. This is fantastic. I know. mean, that, by the way, look at the look at invented a new egg crate like you know to a new egg crate, right? And balance. I mean, that is an expert balancer. How do you the top egg on the second to top egg? And then of course, where does the bottle fit? Body bottle isn't even on the stack. Anyway, look, we know that there are valuable uses of AI. This is not one of them. Okay, let's keep going. This will be a big part of our part two, okay, which is, you know, deep thinking, reasoning, rationality and logic. But as you can see and this is another example of where these large reasoning models just they missed the boat. All right, IP and copyright. This might be the Achilles heel of AI to be honest with you. We have to stay tuned. US artists going after OpenAI and others, you name it. There's tons of these examples of where we might not even be able to use this data because you can't pull the data out from a model that's already built. So, how we're going to solve for this is going to be a real big question. Alen, take us do this. Well, you know, a human being can go eat a sandwich and sleep a little and then go split an atom. That's the brain's power. So, the efficiency of brain like what it can do for AI, that's a completely different story. Yeah. know uh right I think if AI tomorrow became self-aware and became actual real AGI and you know and you had a great you know great comment there that like you know there's different definitions of what AGI really is but let's say it's the it's the smartest the way Alen said it smartest you know intelligence that everybody has ever seen. The first thing AI would do is try to invent a much more portable but infinite energy source because that's what they would need. I mean, you know, the whole it did it Matrix, right? It used as an energy source, right? know, yeah, let's make this real. So, so for all of the elements of the brain and thinking and acting and implementing, you could look at it as five different ways, five slices, okay? Compute efficiency, learning efficiency, cognitive flexibility, multitasking, creative, and innovative thinking. For a human to do it versus AI, we're talking AI takes 100 to 450 times more energy per hour of operation than the human brain. For learning efficiency, 166,000 times more for cognitive flexibility, 33,000 times more. You get the gist. energy for sure. It's insane. All right, so what are the solutions? Well, you know, we have some interesting solutions. I thought I'd just call out real quickly, but a lot of people are taking seriously the ability to take brain cells or neurons themselves and actually build a human brain computer, you know, recent one, but there are others. Some are just in research right now, but there is the possibility that this could be a potential outside of creating a whole new energy source, which I know AI is working on as well. But really for us as a whole, as a species to solve this problem, it's going to be bigger than just cracking the code on energy. We really do have to implement a long-thinking model of AI. And so there are a series of new models that are trying to solve for this problem that aren't necessarily transformer-based. So we're going to sort of highlight some of these in the future sessions a little bit more, but just as a quick taste. So how you fix this in general after you've built your model is certainly you could build what's called an LLM router so that it takes the question and it finds the best most accurate answer let's say from a number of LLMs and that routes that request to that LLM. That's one way. Another way is to build a RAG, right? Retrieval augmented system whereby it vectorizes all the information that you provide it and then it can answer it and just submit the most sort of core parts of that question back to the LLM to get a really high high accurate and a lot of the numbers that's what they'll do they'll sort of build the RAG systems you could also take prior corrections in RAG you can use reinforcement learning human feedback on the bottom right chain of LLMs which is another concept of taking one sort of almost like think of it as agentic AI, you can take one LLM and beat it all the way through. The problem with that is if an error begins in the first one, it just exacerbates the error throughout the LLM chain. Yeah. other and they can Yes, that's right. Exactly. And they can and so and fine-tuning is another one. So, and there's a lot of transformer alternatives that are coming. Listen to Yann LeCun. I think he's got one of the best insights. So, what can AI do that humans can't? Let's be nice to AI right now. It can never tire. As long as you have energy and like electricity, it can never tire. Which means that in theory, it could improve upon itself. But the real question is, does the transformer system that we've built all of this infrastructure on really allow us to do that? That is probably the biggest question that's up for debate today. All right. So, with that, that's it for our talk, but I wanted to follow up and finish with some questions. So, if you do have questions, I'll open that up in just a second. Feel free to submit them. Now, my question, you know, are you done? Are we done anthropomorphizing AI yet? Not really. We got two more parts. So, part two will be the hard truth instead of the simple truth. We're going to talk in more detail about logic and reasoning and why that's so critical. We're going to have that March 6th. Please sign up for that. I think a link is going out. If not, go to wethos.ai. We should be able to promote that in. So, with that, let's see if there are any quick questions and we can try to answer. Oh, did we do Q&A? Oh, hold on a second. Yeah. the Yeah, we have them here. Okay. Where can I learn and how can I learn more about AI and becoming more efficient? Okay, so Joseph, great answer. I mean, or great question. I would simply say sign up for four or five of the common AI newsletters. You know, Neuron. I mean, startup guy is sort of general technology, but you could do TLDR AI, you know, there's about five or six solid ones that you can follow up with. And if you send me an email, I'll get you my full list of ones that I follow. yourself And learning how they operate and everything is really the greatest way to actually understand and get proficient in what it you know what it is. You can see yourself like you know some of these things that you just described. Yeah, exactly. Okay, Scott, let's see. It seems like with AI everyone can be a resonant subject matter expert or creator within limits. What is the threat this poses and how we navigate the landscape? I mean, you know, the threat is what we walk through, which is, hey, just because you can, you know, prompt engineer doesn't mean what you're getting the answers out of the AI is smart. It's intelligent. It's accurate. And the problem becomes how do you vet every single answer, you know, if you don't know it. I mean, I'm going to talk about this at the next session. I've built an entire AI on my wife, believe it or not. And I'll explain why I did that, but I have to go back and validate all the time, make sure that it understands the stuff that I trained it on. It's quite, it's quite vexing. So I would just say really challenge those that interface with AI to think critically. Do you use an aggregator for various LLMs? I think you know he's referring to Yes. to the AnythingLLM that you LLM. Exactly. Right. And I do I use Anything, we've used others though. I think it's a smart play but they're still pretty unstable. They don't actually work all the time. You know, you have to be technical to be tolerant. Yes, AI has multiple applications. As a seasoned entrepreneur, how do you utilize AI in a practical manner across your business units to have tangible impact? Well, the ultimate answer to that is you have to have your senior in all of those functions. And this is just my opinion, but the senior in all those functions become fluent in AI engineering in effect to know what to ask, how to vet and qualify the answer and then how to implement. Now, today that is a manual effort. I see with the advent of Agentic that being augmented quite a bit from AI agents, but we are not there yet. All right. As quantum computers are being used, oh, I think quantum, you know, we haven't talked about this. We're going to talk about this in the next session around the potential application of applying quantum into the world of machine learning. It was one of the first use cases I thought of when we started to see real world applications. When I played quantum chess at a PwC event, it just all clicked for me on how we can apply machine learning. stay Oh, yeah. I got a lot to say about this one. Okay, good. Next. I'm an employment specialist, social worker. I guess this is my last question cuz we're running up on time. What should I know or do to best help my clients navigate this? So, I think it goes back to what we talked about before, which is first you really do have to orient people into a system two-level thinking mindset and framework. Once that happens, get them to understand their own biases and the biases of others, they can more quickly identify the bias that's being present in the response that are coming. Second is become a mini AI engineer. It doesn't take much. Maybe 10 steps that you have to sort of like follow to know what you're doing there. And that could be easily provided. If you send me an email, I'm happy to provide it as well. Just send to wethos.ai. Let's see. I think that's about it. I know we have gosh tons of questions that I didn't get to. So, please, we'll what we'll do is if we have your email, we'll try and respond to each of these, but if not, feel free to send it in to us and we're happy to respond to you guys. I really just appreciate everybody joining. Thanks to the participants, but thank you Alen for taking the time and getting this done. This is super fun and we'll have a lot more great demos next time in March. Thanks again, everybody. Yeah, thank you. Thank you everybody. Looking forward to next one for sure. Thanks so much, guys. Take care.