What AI Won’t Tell You When It’s Wrong | Jonathan Schaeffer
In this episode of An Hour of Innovation podcast, Vit Lyoshin speaks with Jonathan Schaeffer, founder and CEO of Synsira and creator of Kind, about why today’s AI tools can be useful but still unsafe to trust blindly. Jonathan explains why large language models such as ChatGPT, Gemini, and Claude should not be treated as intelligent authorities, why the word “hallucination” makes fabricated answers sound too harmless, and why a polished AI response can still contain errors users may never catch.
Jonathan shares how ChatGPT once generated a biography about him with dozens of mistakes, then appeared to improve later through what he calls band-aids rather than a deeper solution to the trust problem. He explains why training on broad internet data can produce strange or unsupported answers, and why curated datasets can reduce fabrications by grounding responses in information the user controls. The conversation also covers source citations, verification, and why trustworthy AI should be able to show where each fact came from.
Vit and Jonathan discuss Kind, Synsira’s local AI product that analyzes documents, presentations, images, videos, and personal knowledge collections on the user’s computer. Instead of moving private files to cloud AI providers, Kind moves AI to the desktop so answers can come from local data and cite original documents. Jonathan explains why this matters for books, intellectual property, financial records, company strategy, customer-confidential documents, and other sensitive files.
The episode also explores why 95% or 995 accuracy may still be unacceptable when users cannot tell which answer is wrong, why friendly chatbots can encourage misplaced trust, and why Jonathan prefers “augmented intelligence” over artificial intelligence as a product philosophy. For AI users, founders, product leaders, and knowledge workers, this conversation offers a practical framework for using AI with better judgment, boundaries, and more demand for proof.
Jonathan Schaeffer is an AI researcher, founder and CEO of Synsira, and creator of Kind. He spent 40 years at the University of Alberta, where he worked in artificial intelligence, research, teaching, and public speaking before leaving the university to build a local AI product focused on privacy, verifiable answers, and user-controlled data. His background also includes pioneering work in game-playing AI, including Chinook and computer checkers, and recognition as an AAAI Fellow and Royal Society of Canada Fellow.
You’ll Learn
- Why calling AI tools “intelligent” can create dangerous misunderstandings
- Why hallucinations are better understood as fabrications, errors, or lies
- How ChatGPT changed public expectations after November 30, 2022
- Why students and everyday users can overtrust confident AI answers
- What Jonathan’s incorrect AI-generated biography revealed about band-aid fixes
- How curated datasets and source citations can make AI answers easier to verify
- Why local AI can analyze personal files without moving them to cloud AI providers
- What Kind does for documents, presentations, images, videos, and personal knowledge collections
- Why cloud AI creates privacy risk for books, intellectual property, company strategy, customer documents, and personal data
- Why 95 or 99 percent accuracy may still be unacceptable when users cannot tell which answers are wrong
- Why Jonathan prefers “augmented intelligence” over artificial intelligence as a product philosophy
- What basic AI education every user should have before trusting these tools
Timestamps
00:00 Introduction
01:18 Understanding AI Misconceptions
04:55 The Trust Dilemma in AI
13:19 Hallucinations vs Accuracy in AI
20:20 Introducing Kind: A Local AI Solutions
29:39 Curating AI Models for Optimal Performance
34:14 Understanding AI’s Practical Applications
37:20 AI as Augmented Intelligence
42:08 The Limitations of AI and Public Perception
48:06 Machine vs Human Learnings
52:58 Educating Users on AI Interactions
58:00 Innovation Q&A
Connect with Jonathan
- Website: https://kind.synsira.com/
- LinkedIn: https://www.linkedin.com/in/jonathan-schaeffer-phd-frsc-aaai-fellow-3318015/
- Other: https://en.wikipedia.org/wiki/Jonathan_Schaeffer
Connect with Vit
- Substack: https://anhourofinnovation.substack.com/
- LinkedIn: https://www.linkedin.com/in/vit-lyoshin/
- X: https://x.com/vitlyoshin
- Website: https://vitlyoshin.com/
Support the Podcast
If you enjoy the podcast, you can support it by exploring the tools below.
These are affiliate links, meaning the show earns a small commission at no extra cost to you.
- MeetGeek: Record, transcribe, summarize, and share insights from every meeting.
- Amplemarket: Step into the future of sales: Human + AI. Empower reps, uncover opportunities, and grow revenue.
- Databox: Turn business performance data into clear answers your team can understand, explain, and act on – instantly.
For inquiries about sponsoring An Hour of Innovation, email iris@anhourofinnovation.com
Episode References
ChatGPT
https://chatgpt.com/
ChatGPT is OpenAI’s conversational AI product mentioned throughout the discussion of hallucinations and trust.
Google Gemini
https://gemini.google.com/
Gemini is Google’s AI assistant mentioned as one of the major cloud AI tools users compare with ChatGPT and Claude.
Claude
https://claude.ai/
Claude is Anthropic’s AI assistant mentioned in the conversation about choosing among cloud AI tools.
Meta Llama
https://developer.meta.com/ai/models/llama-4/
Llama is Meta’s family of open AI models mentioned among common AI tools users may encounter.
Large language model
https://en.wikipedia.org/wiki/Large_language_model
Large language models are the AI models Jonathan discusses when explaining probabilistic answers, fabrications, and hallucinations.
AI hallucination
https://en.wikipedia.org/iki/Hallucination_(artificial_intelligence)
AI hallucination is the term Jonathan criticizes as a softened label for fabricated or erroneous AI answers.
Wikipedia
https://www.wikipedia.org/
Wikipedia is mentioned as a relatively reliable source that AI systems may use or retrieve for biography-like answers.
Encyclopaedia Britannica
https://www.britannica.com/
Encyclopaedia Britannica is mentioned as an authoritative source AI systems may use for biographical information.
IMDb
https://www.imdb.com/
IMDb is mentioned as a source AI systems may use for biographies of actors and directors.
Google Search
https://www.google.com/search/howsearchworks/
Google is referenced when Jonathan contrasts web-scale indexing with the smaller scale of a user’s personal data.
University of Alberta
https://www.ualberta.ca/
The University of Alberta is where Jonathan says he spent 40 years before leaving to start his company.
OpenAI
https://openai.com/
OpenAI is mentioned in relation to fast-changing chatbot leaderboards and the broader AI tool market.
Star Trek
https://www.startrek.com/
Star Trek is mentioned when Jonathan describes the factual, business-like computer as an inspiration for augmented intelligence.
Vit Lyoshin (00:01.09)
Welcome Jonathan. Thank you for your time today.
Jonathan Schaeffer (00:03.9)
It's a pleasure. I'm delighted to be here.
Vit Lyoshin (00:06.606)
Sure. so I wanna start this conversation. people use AI today and it gives pretty intelligent answers. but as I understand your argument is that you cannot always trust AI tools. what are some of the like biggest dangerous misconceptions that you think people have about this?
Jonathan Schaeffer (00:30.75)
Actually, you said the very first one. You describe these programs as intelligent. I disagree. We tend to we tend to anthropomorphize them. We think of them as human or human-like. And yet, if you look at what they're doing, if you if I were to prep pretend I was a professor and explained to you the algorithm, you would say, that is so incredibly stupid.
Vit Lyoshin (00:38.392)
Yeah.
Jonathan Schaeffer (00:57.844)
There's no intelligence there. It's dumb. But it works. Okay? The mantra in the field of artificial intelligence is you do whatever you want. It's the result that counts. And in the case of these large language models, these programs like Chat GPT and Gemini and Claude, there is no intelligence. So
That's the first thing that we need to understand. Because that explains a lot when you start talking about the answers that you get back from these programs. We get these hallucinations, which again, there's another anthropomorphization. Hallucination. Isn't that cute? Let's be candid. What did it do? It lied. It fabricated. It's an error. Why don't we talk about ChatGPT and say,
What's its error rate? no, it had a hallucination. It's like we're almost apologizing that it okay, maybe it it was on drugs or it wasn't feeling well. It just had a hallucination. But again, when you think it when you look at the algorithms, because there's no intelligence, it doesn't really understand what you're asking. And therefore, you'll get responses that.
Have no relationship to what you actually asked. And although this is very simplistic, at the core of the algorithm is a coin flipping mechanism. And sometimes it flips the coin in the right way and it looks brilliant. Okay, I'm anthropomorphizing, but from your point of view, it may look brilliant in its answer. And sometimes you flip it and the coin goes the other way and it looks incredibly
Vit Lyoshin (02:43.021)
Mm-hmm.
Jonathan Schaeffer (02:48.719)
silly and stupid. This this is sorta I guess my what I'd like to believe and I hope that I'm right is that there are fundam I know that there are fundamental flaws in what we're doing today and these problems aren't gonna w go away. Hallucinations are there, they will always be part of the algorithm. It's it's never going to be solved.
using the ways that we currently build these programs. But I hope what we have what's going to happen is that this is a stepping stone to a better technology, one which we've learned a lot from the current models. And this new hypothetical technology addresses those problems and gives us a much better, more accurate, safer, private, whatever, insightful experience.
Vit Lyoshin (03:42.336)
Mm-hmm. Yeah. Okay, great. And I know you have a long history of working and researching and building with artificial intelligence in general, but when if first you started realizing that there is a problem with trust with these tools?
Jonathan Schaeffer (04:01.775)
so I'd seen and played with earlier versions of Chat GPT. But then on November 30th, 2022, my world changed. Now you're sitting there thinking, November 30th, 2022. What happened that date? Well, that's the date that the new Chat GPT came out. And it rocked my world because I've been in AI for a long, long time. And it did things that I never thought.
I would see anytime soon. And I was stunned because it was so much better than what we'd seen before. The answers were better, the capabilities. And of course, this was a big media event. And ever since that date, the world has gone AI crazy, and we've seen, I don't know, countless trillions of dollars going into AI or AI infrastructure. But very quickly, problems started surfacing. The first
Vit Lyoshin (04:54.136)
Mm-hmm.
Jonathan Schaeffer (04:59.855)
of course was the hallucinations, the errors. And in the early days they were quite frequent and they were they were laughable. It was it was funny and it was amazing. But later on at the university, it became quite unsettling because students started using it. And because this is an AI program and it's the latest algorithms and it's
Vit Lyoshin (05:09.176)
Mm-hmm.
Jonathan Schaeffer (05:28.467)
Been trained on all the world's data. They believed what it said. It like there was this trust relationship that I'm just a student, and here is this super intelligent program that's going to give me all the answers. And clearly that's wrong. But the trust level between the student and the program was such that.
They weren't questioning the results. They believed it just because this is like amazing. And that was unsettling because as a researcher, I know how the algorithm works and I know that it can't be trusted to be accurate. What we've seen since then, and this will sound rather cynical, and the AI companies probably don't like the word I use, but
They've applied a lot of band-aids. So the errors are still there, but they put on all these extra patches, fixes to cover scenarios that people point out that are embarrassing. So, just as a personal example, when ChatGPT first came out, I said, Hey, who was Jonathan Schaefer? And it gave a bio and it was really interesting. And I showed it to somebody and they said, Wow, what an incredible career you've had. Look at what you've done.
Look at all those awards. That's amazing. And then I pointed out 21 errors in a page and a half of text. And I gave a live demo on this of this program of ChatGPT and asked if my bio about a year and a half later, and it had 23 errors. And then about a year later, or less than that, I gave another demo and it was almost perfect. What? What?
Well, it seems a lot of people were asking for their bios. So what does they do? They added a band-aid. if this looks like a biography, go to Wikipedia. And if you can find the person, use the Wikipedia. and if it's not there, look in an authoritative source like the Encyclopedia Britannica. Actually, you and I are of the generation we know what the Encyclopedia Britannica is. I used that once in my classroom and the students said, What are you talking about? They had no idea what that was. But
Vit Lyoshin (07:30.274)
Mm-hmm.
Vit Lyoshin (07:43.0)
Yeah.
Vit Lyoshin (07:48.052)
Ha ha ha.
Jonathan Schaeffer (07:50.12)
And of course, you can go to the internet movie database to get biographies of actors and directors or history databases, and it goes out and searches to see if it can find a biography, and then it ingests it and spews it out. And guess what? It's almost perfect. And so that's an example of a band-aid. And there's all sorts of other band-aid, hundreds, thousands of band-aids that are applied. And the errors are there, they're disguised, right? It's like,
Vit Lyoshin (08:05.518)
Yeah.
Mm-hmm.
Jonathan Schaeffer (08:19.049)
To use a graphic image, somebody got really seriously injured in a car crash. it look at look at all the cuts, look at all the bruises, the broken bones, this, that, and the other thing. We're gonna take a band-aid and we're gonna put a band-aid on this cut and we'll put another band-aid on that cut. And we'll wipe the blood away here. And the person looks better. They're not bleed they're not bleeding as much, it doesn't look so horrible. But you haven't solved the real problem. The real problem is, you know, maybe.
maybe their lungs have collapsed. You know, all these cosmetic changes don't fundamentally change the nature that the patient is sick, really sick. And maybe that's a bit too dramatic for these programs, but the fundamentals, what's inside these programs, still has problems. The band-aids hide it, they don't solve it.
Vit Lyoshin (08:55.852)
Yeah.
Vit Lyoshin (09:11.07)
Mm-hmm. It reminds me a story that when I first came to my computer science class long time ago in the university, and and professor said basically, look around you, who do you think is the soup the stupidest here in this room? And we were like, What are you talking about? How you don't even know us? And he said, It's computers in front of you. So those machines are like, you know, dumb iron things. And now we we pretty much
saying the s saying the same thing about AI because we realizing more and more with our experience that they are not that smart and I use AI every day and I see issues left and right myself and I sometimes spend more time messing with it and correcting it rather than just doing from scratch.
Jonathan Schaeffer (09:58.88)
but you are the smart one because you recognize that it's a problem. And when you get an answer back, you do the homework to validate whether the answer is correct. I don't want to be embarrassed. I write blog posts, I write books, I give public speaking. I don't trust these programs to give me the right answer, which means I spend the time to verify the facts before I use them because I don't want to be embarrassed in public. But
Vit Lyoshin (10:14.925)
Uh-huh.
Jonathan Schaeffer (10:27.773)
What's surprising about these programs is they can be incredibly useful, but I still have to invest time to make sure that I'm getting back the information that I need and that it's correct.
Vit Lyoshin (10:40.12)
Yeah. Yeah, and I think this is critical, especially for students, like you mentioned. It's amazing how they believe that right from like it's like believing a stranger on the street, whoever says something. It's it's I don't know. I I've been, I guess, raised in a different manner. My parents always told me don't trust strangers by default. So yeah, we'll we'll see how it goes.
Jonathan Schaeffer (11:05.491)
But you know what? this is one of the things that is very disturbing about some of these chat bot programs. And again, the word chat is an anthropomorphization. But
Vit Lyoshin (11:14.211)
Yeah.
Jonathan Schaeffer (11:19.113)
They become more friendly. They're they they talk to you in a in a way that tries to build a relationship. It gives you an answer and it says, would you like some more information? Or how can I be a further help? Would you like me to explain this, or do you want me to do that? So they're trying to create this symbiotic relationship between you and the computer program. And of course, there have been
Many stories of with horrible outcomes where people have put trust in these programs and the program has given them bad advice and they've acted on that advice, and sometimes bad things happen. It's it's actually a a dangerous situation when we're trusting a technology that we know has some flaws in it.
Vit Lyoshin (12:08.13)
Yeah. So given LLM's probabilistic nature, is there a way to make them like give you correct answers all the time?
Jonathan Schaeffer (12:21.385)
So that's a good question and I wanna differentiate be between correct and hallucinate. And people sort of put the two words together, but I I view them as differently.
Vit Lyoshin (12:28.525)
Mm-hmm.
Jonathan Schaeffer (12:37.201)
If you go to an LLM and ask it a question, it's been trained on the world's data. Everything it it can find. I mean it's it's insatiable. It wants more and more data. And they've scraped as much data as they can get from the internet, right?
All of the Wikipedia is inside, all of the the internet movie date databases is is in there, every commercial site that's publicly available is in there, every book that's online, every pay research paper that's on, it's all in there. But they still want more because the results show more data leads to gradually better gradually better performance. So
Sorry, I just lost my my thread on this. sorry, could you just repeat the question? Right, sorry, thank you. Right. So they've got all this information and remember they don't understand the question. So sometimes you get these these f these hallucinations, and they're funny in particular because it's pulled some data that it's been trained on.
Vit Lyoshin (13:34.136)
can we make those LLMs give you correct yeah, mm-hmm.
Jonathan Schaeffer (13:53.95)
And that data could be really good, like the Wikipedia is pretty darn good. It's excellent, in fact. But there's other sources that it's trained on which are suspect. And so it pulls them in and it becomes part of your answer and you get these strange situations which we call hallucinations. But there is a solution.
Jonathan Schaeffer (14:14.683)
If you have a data set that you've curated, it could be a personal data set, it could be your company's data, it could be all of just Wikipedia. You can build these LLMs, large language models, or something smaller, an LM, like a language model, and have it trained only on the data that you want it to be trained on. In other words, you curate the data.
And you say, this is the only data that I want it trained on. You've controlled it. And of course, if I'm gonna do that, I'm going to ensure that in my data set is only data that I think is useful, right? It's stuff that's valuable, informative, that I've believe is is worth having. And I'm not gonna put in all the junk. And then if you've filtered out all the junk.
These language models will only extract answers from your data, which means they're not going to hallucinate. Not in the sense that what you'll get from, say, a Chat GPT on occasion. However, that doesn't mean the answers are accurate or correct. They're just not fabrications. The answers will only come from your data, but they still might be wrong because the program may not have act correctly interpreted your question.
misunderstood it, perhaps not giving you enough of an answer, or maybe it's it's overdone it. But that's that's not a error per se. It's more of insufficient information in the data set or insufficient information in the question that you ask, the query that you make. So you can create an environment where they don't hallucinate and fabricate answers, but it's still
the reality that you're going to ask questions and perhaps not get back quite the right answer that you expect. But that's the same with you and me. You can ask me questions and I can might misinterpret what those questions are and give you back something that's not what you expect. And then what do you do? you say, yeah, that's not what I meant. And then you give me the question with a bit more detail. And of course, if you do that with one of these large language models or language models, you get better results.
Vit Lyoshin (16:29.231)
Yeah.
Vit Lyoshin (16:36.895)
Mm-hmm. Yeah, yeah, that makes sense. Okay. Because I was thinking a little bit in a different direction, like maybe give some sort of citations or like verification badges like we've seen in like social c in social media, for example. So those type of or maybe where which sources came from, so you can go verify kind of like an academic paper where you cite where you get your sources from. but
How realistic is that? I'm not sure.
Jonathan Schaeffer (17:06.921)
Course, it it's there are things that you can do with your data that companies like Google can't do. Because your data, doesn't matter what your data is, it's small. As opposed to Google, which is trying to, you know, map or understand the entire internet. I mean, what do you have on your laptop of data? Some gigabytes, tens, hundreds, gigabytes? Maybe you have a hundred gigabytes of data there.
Vit Lyoshin (17:21.647)
Mm-hmm.
Jonathan Schaeffer (17:35.966)
That's small. What's Google trying to index? Petabytes or Yada bytes or who knows what? Just massive amounts of data, right? And if you have a small system like your data, your personal data, data that you've created, data like, you know, my financial records or my summary of my investments, or you're writing a book and this is all the data to do with your book, or you're you've got a passion and you're collecting.
Vit Lyoshin (17:45.326)
Mm-hmm.
Jonathan Schaeffer (18:06.335)
stamps and you've got you know information about all your stamps. You've curated the data, you've s chosen the data that's relevant and that's useful to you. And then you can train these programs so that every fact is mapped back to the original source. And then when you ask questions, you can get the source for every single fact that's in your answer.
Because it all comes from your data. And you can do that because your data is small and it's easy to maintain all of that information. In other words,
You can do things for your data that Google can't do for the world's data, or they could do, but you know, the cost would be just enormous to do that. And so there are advantages to thinking about large language models or language models that operate not on the world's data because you have no idea what's in that data. Nobody's told you what's in that data, as opposed to doing it on your data where you know exactly, sorry, not where.
Vit Lyoshin (18:59.246)
Mm-hmm.
Jonathan Schaeffer (19:17.395)
You know what's exactly the data that's being it's being used to train the answers.
Vit Lyoshin (19:21.764)
Mm-hmm.
Yeah, yeah. Okay. I think this is a good time to switch to talk to about what product you build in, which I think is called Kine, right? can you tell us a little bit what it is, and how it operates and things like that.
Jonathan Schaeffer (19:39.391)
Sure, so you already touched on part of the motivation. I just got
upset with these large language models and it it was
Jonathan Schaeffer (19:54.108)
It was several things. It was clearly the hallucinations. Because as a scientist, I want I want accurate answers. I don't want to have to question everything. I also was annoyed at the attempts to anthropomi anthropomorphize these chat programs. They became my friends and they treated me like a friend, very chatty, very friendly. And no, you're not my friend.
You're not intelligence. This is just a a programming trick. It's a gimmick to try and make it look friendly. It's like you know, every time you log into one of these computers and says, hello, well, a friend will say hello to you, but a computer saying hello is nothing. It's just printing texts on the on the screen. There's no emotion, there's no feeling there. So those two things were annoyed me. And then of course.
I got further annoyed when I understood how these programs were hungry for data, all sorts of data. And when you've exhausted the internet, they'd love to get access access to all of your personal data. Your data is your data is gold, and it's not just your data files. That's the most visible data. It's everything you do on the internet.
What what websites do you go to? How long do you spend on a web page? What do you read? What do you purchase? Who do you send messages to? That's all data. It shows human behavior. the text shows grammar and how to use it use it. there's facts in there. I mean, these programs want that data, and again, more data translates into incrementally best more performance. And then of course, there's the issue of the environment.
I am really stunned at how much money is being invested in data centers. When you think of how many globally, how many trillions of dollars are going into data centers? How many of the world's problems of poverty, for example, could be solved if all that money was just applied to making this world a better place than to build data centers? Data centers that use a lot of water and energy.
Jonathan Schaeffer (22:22.799)
and environmentally unfriendly. I mean, okay, I'm sure I sound like I'm a real green environmentalist, but there are a lot of problems in this world that we could solve with the vast amounts of money like that's going into data to AI data center infrastructure. So that was the motivation and
I've been at the University of Alberta for 40 years. And after 40 years, I've had a great career. I've published lots of papers, worked with great students, taught amazing students. But it was time for a change. And it wasn't enough to be able to talk the talk and say, hey, I think there's a better way. I had to walk the walk. So I left the university and I've started a company and we produced our first product. It's it's been out for six months, and it's called Kind.
which originally stood for Knowledge Index and now stands for Knowledge in Depth. So it's sort of an acronym. And what it does is try and address the problems that I identified. So think of it this way: you've got your own Chat GPT on your desktop. It's more than that, but that's an easy way to think about it. The idea is
Instead of moving your data to the cloud to get access to high performance AI, we move the AI from the cloud onto your desktop so that you can get high performance AI for your files. So if you want to ask questions about the world, yes, you have to go to ChatGPT or Google or Cloud. But if you want to ask questions and understand your data, your financial data, your family history,
The book that you're writing, your stamp collection, stuff like that. What this program does is analyze the data on your computer and allow you to s ask questions and do things using only your data. So if you ask a question, the answers only come from your data, and all the information is cited back to the original document that's on your computer, right?
Jonathan Schaeffer (24:39.911)
And so the fabrication problems go away. It it's it's highly accurate. Sometimes it'll make mistakes because it misunderstands you or doesn't give enough detail. But that's a different kind of problem. So it's 100% private. Your data never goes to Google or any of these other large language model providers for training. So you've got the privacy that you want.
Vit Lyoshin (24:46.084)
Mm-hmm.
Jonathan Schaeffer (25:07.443)
You're not sending queries over the internet, you're not paying token charges, you're not paying $20 or $30 a month to get a an internet large language model account. but one of the things that's important to me again is the environmental movement and their mantra is think globally but act locally. And in some small way, I like I like kind to be part of that movement because
I'm thinking globally that you want access to high performance AI, but you don't want the environmental footprint. So let's put it on your desktop. And now you're not using these data centers. If we all move started using local AI, we wouldn't need these data centers. And if you think about local AI, it makes sense.
I don't know what computer you're using right now, but I got a computer in front of me. It's a Mac. It's got 18 cores in it. It's got 24 gigabytes of memory. It's a powerful computer, right? And for most people, that computer, you've already paid for it. It's a sunk cost. It's sitting idle. Is it busy during weekends? Is it during busy during the evening or all night long?
Is it busy on holidays? Even during the day, my computer's largely idle. I've got so much computing resources right in front of me. Why not use it for the AI? There are things that we can do in kind that Google can't do for you. And the reason is because we've got your CPU. We can use it to do all sorts of deep AI analysis on your data for you. And it doesn't cost you anything because you've already bought the computer.
You could do these same same things using agents with Claude and ChatGPT and Gemini, but you'd have to take your data and move it into the cloud and then pay the price for having all these agents to do the analysis and you'd be facing token charges. Why do that? Just do it on your desktop. You're not paying for any infrastructure, you're not paying for any tokens.
Vit Lyoshin (26:52.675)
Mm-hmm.
Jonathan Schaeffer (27:18.939)
So thinking globally and you're acting locally, you're not feeding the data center craze. So for all of those reasons, that's why we did it. I believe, perhaps naively, but I believe that we can we can help course correct where AI is going and make it about the individual.
Vit Lyoshin (27:26.115)
Mm-hmm.
Jonathan Schaeffer (27:43.314)
And I think that most individuals have a lot of data, useful data, important data, personal data on their own computers that is untapped. And why not? I'll give you a perfect example. I have 40 years of presentations, public speaking, lectures I've given in courses, lectures I've given at conferences, right? Hundreds of presentations.
Vit Lyoshin (28:06.831)
Mm-hmm.
Jonathan Schaeffer (28:09.191)
I just say, hey, kind, I'm going to create what we call a collection. Here are all my presentations over the last 40 years. Analyze them. And it does. And now I can ask questions. I can find out, hey, I gave that talk on that particular day. Or I can say, I think I once talked about something and it'll find it for me and tell me back. Or I'm going to give a new talk and it will tell me, you know, the pieces from previous presentations that I could use to create my new presentation.
It's all coming from my data, right? And I find that exciting.
Vit Lyoshin (28:39.948)
Uh-huh.
No, that's great. how does it work on the back end? Is it a like a pre trained local model or something? Or
Jonathan Schaeffer (28:50.335)
So yeah, good question. So first of all, we don't do training on models. So what happens is when you you you get kind, it downloads basically two things, the executable program, and it uses an open source model. we choose the open source model based on models that are we we at Sincera believe are
Vit Lyoshin (28:58.318)
Mm-hmm.
Jonathan Schaeffer (29:20.039)
you know, highest quality. they're not biased, they're fair. and based on the resources on your computer, you might get a big model or a small model, depending upon how powerful your computer is, how much memory, how much disk. And then we have our AI, which is which is proprietary, that does use this model to help us analyze your data. So for example, a big
component of what these models do for us is help us understand videos and images. So we didn't write the code to do video analysis. We use these models to take pictures from your documents or maybe your family photos and analyze them for you. Or maybe you put in your favorite videos and we do the analysis of the speech and the pictures inside, sorry, the images inside or the frames inside your videos. And we we use those models for it. And we also use the models for grammar.
Vit Lyoshin (29:52.568)
Mm-hmm.
Jonathan Schaeffer (30:16.147)
Because once we've ask a once you ask a question, we need to be able to take the facts that we find and stitch them together into a readable answer for you. And those models contain that. So the point here is that you use our code and the tools that we've built. And behind it, they're also using these models, public domain, open source models to help generate the answers.
But there's one subtlety in here that I think is really important and it excites me, and maybe it doesn't excite everybody, but it's stunning how fast the world of AI is advancing. it's just it's unprecedented in human history. And you know, if you go to the the leaderboards that compare these the chatbot programs and can say, Hey, look, today OpenAI's got the best chatbot according to this metric. And
Vit Lyoshin (31:00.142)
Uh-huh.
Jonathan Schaeffer (31:11.623)
Two weeks ago it was Gemini and a month ago it was Claude and tomorrow, no, there'll be a new announcement, and it's it's who knows what is now the best one. It's changing so quickly. And this all came out of a user who said, I don't know what I want to use AI, should I be using Claude? I hear all these things about Claude, but I grew up using ChatGPT and I trust Google and which one should I choose? I I don't want to have three accounts, right?
Vit Lyoshin (31:40.108)
Yeah.
Jonathan Schaeffer (31:41.204)
And what we do with our product is those open source models, we do the your work, we decide which is best. We you're curating your data, you have your data sets, you've chosen what's best. Sincyra does the same thing with the AI. It's our job to understand what's happening in the world and choose the best model. So, for example, if you download
Kind local pro today. Inside, it uses an open source model from Google. But if you were using our testing our product a month ago, it was using a different model. And in that time period, we said, you know what?
This this new model from Google, it's better than the one we were using before. So we swap out the old and we swap in the new. And we're going to be releasing a new version of our program in a couple of weeks. we may swap the model again just because we think, because the few the advances are happening weekly, that there's even a better model out there. Maybe it's it's smaller or more powerful or it's
more free of bias or whatever, and we may swap in a new model, which means you don't have to worry about it. It's not your job to understand what's happening in the AI field. It's our job to make sure that you get the highest quality AI experience.
Vit Lyoshin (33:04.889)
Mm-hmm.
Vit Lyoshin (33:17.615)
Yeah, and I mean for most everyday tasks, all of those models are fine. It's not like we're doing high precision calculations of some sort like you know, I don't know, astrophysics math or whatever. So so
Jonathan Schaeffer (33:33.096)
And and that's true. You know, one of the things that's really important for people to understand is that when you hear a lot of things about these models and the incredible things that they're doing, it's on the high end. The it's these are capabilities that the average person doesn't need. I mean to have these models you know, figure out ways to you know,
break into websites, which is something that's been happening recently, or using them to discover new medical treatments, or use them to unwell, to use your example, to help them understand new properties of physics, you know, okay, that's research. Maybe that's for the big companies. It's not for you and for me. For most things that you and I do, which is reading, writing, asking questions, finding information,
Jonathan Schaeffer (34:30.569)
The capabilities right now are pretty darn good. And an incremental improvement of a percent or two or five percent really makes no difference. And if you're, you know, if you're paying $20 or $30 US a month for an internet provider to do those kind of things, if you you know, if you're paying $30 a month to Claude and all you're using for Claude is to give it your data, which it might use to train on, and ask simple questions and
you're actually not spending your money wisely because a lot of that can just be done on your desktop, on your computer, with a one time investment in software and not paying token charges and not paying monthly fees.
Vit Lyoshin (35:15.449)
So so people will buy product one time, just like for example, like Windows operating system and then get updates once in a while.
Jonathan Schaeffer (35:23.581)
Right. So the way it works is you purchase it once, seventy nine dollars in the first year, and that you get all sorts of updates, which are bulk fixes, and upgrades, which are new features and new functional new functionality and new models, perhaps, et cetera. And then every year there's just a thirty nine dollar renewal fee to keep going.
Vit Lyoshin (35:48.686)
just for support basically. Yeah, okay. Yeah, that's pretty cheap. That's a lot cheaper than most of these tools.
Jonathan Schaeffer (35:50.494)
Yeah.
Jonathan Schaeffer (35:55.934)
Well, and that's that's that's part of our argument. And yes, you're not gonna use our models to do, you know, medical research, we agree. But we're targeting consumers because there are a lot of consumers out there that
Vit Lyoshin (35:59.471)
Yeah.
Jonathan Schaeffer (36:13.287)
really could benefit from AI. One I know this will sound s silly, but
Jonathan Schaeffer (36:23.707)
I'm old enough, and I have the gray hair, I guess, to prove it, but I'm old enough to admit that I saw Star Trek when it first came out. Okay, I was young and perhaps definitely did not fully understand it, but I've watched the original series over and over again. And when I got into AI, the computer in Star Trek was my role model for AI. And you had Captain Kirk and you had Spock.
And they would ask questions to the computer. And the computer was all business. You would ask it a question, it would go off and think. It may answer right off the bat, or it may go off and think for a few minutes, or whatever, and come back with an answer. And it would give you the answer or its best estimate of what the answer was. And if it didn't know, it said it didn't know. And in my mind, that's what AI should be. It's not my friend, it's not something I want to chat with.
Vit Lyoshin (37:05.23)
Mm-hmm.
Jonathan Schaeffer (37:21.851)
It's a resource that I want to use to help me be more productive. The way I like to think about the computer in Star Trek and the computer that we're building with, sorry, the program we're building with kind is that it's a different kind of AI. We think of AI as artificial intelligence, but I like to think about it as augmented intelligence. I've got the intelligence, and I want to use this program to augment, help me.
Vit Lyoshin (37:27.502)
Uh-huh.
Jonathan Schaeffer (37:51.162)
I look at kind and saying, hey, I was a professor. This is like my graduate student. I'm gonna ask my graduate student something to do, and it gives me back the answer. And if the student doesn't know, he admits it. He says, I don't know. There's no bullshit. sorry. There's no fabrications. right? And there's no personality there. Kind has no personality. We're not trying to be your friend. We're trying to help you. And I think that's the best way to think about AI as a tool that helps you.
Vit Lyoshin (37:51.299)
Mm-hmm.
Vit Lyoshin (38:05.866)
Ha ha
Jonathan Schaeffer (38:21.299)
Be more productive or improve whatever you're doing or add value to your life. It's not
your friend. It's not
It's not what AI, it's it's not part of the direction that AI is going today. I think we want to think of it as a tool and not our colleague or our friend, something that we should be talking to. We want something that helps us by giving us fact-based answers. It doesn't lie to us, it doesn't fabricate answers, and it has to say, I don't know, when it doesn't have the information.
Vit Lyoshin (39:01.294)
Uh-huh.
Vit Lyoshin (39:07.203)
Yeah, I I mean people are encouraged to do that and there is no shame in something not knowing something, right? we do all the time, but we expect AI and it probably it feels better than it make up stuff and tell us something rather than being like just blank or I don't know, just go do yourself. We we don't like that because we paid money for it. It gotta give us something.
Jonathan Schaeffer (39:30.846)
Right. Well, you actually you just anthropomorphic, you said it quote, it feels better, unquote. And if that applied to the software, it doesn't have any feeling, right? So it you know, if i i they they tr these programs try to give you an answer, and that answer, whether it's baloney or whether it's gold, they don't understand, they don't appreciate, they don't know, they don't know that they're actually generating work for you by giving you fictitious answers. They're just
Vit Lyoshin (39:39.489)
Yeah, yeah.
Vit Lyoshin (39:46.637)
Yeah.
Jonathan Schaeffer (40:01.183)
Trying, they're just doing what they've been told to do, which is flip enough coins, generate the text, and you judge the quality of the answer. But it's the anthropomorphization that's a real killer to me. when you you know, knowledgeable people like you, myself, and many others, we get it. But having given so many public talks,
Vit Lyoshin (40:15.97)
Yeah.
Jonathan Schaeffer (40:27.849)
To the average person on the street who doesn't understand this technology, they don't get it. They they still think that this is these are super intelligent programs. And it doesn't help that, you know, in the media we're talking about the day that super intelligence comes, or AGI, as it's called, artificial general intelligence, because all that does is cement in people's minds that these programs are smarter than them and can do.
Vit Lyoshin (40:47.715)
Yeah.
Jonathan Schaeffer (40:56.913)
Amazing things. The human brain is incredible. We can do amazing things. And using these programs to help us, we can be better. We can be more productive. We can be more creative. And that's what I want to see.
Vit Lyoshin (41:12.161)
Yeah, no, that's great. And then we hear stories when companies actually employing a lot of these AI agents, replacing jobs, real jobs, right? And and or trying to substitute a lot of them, like software vibe coding type things. many people writing software code. And actually, I talked with a few people who are now saying you should not be re deploying your code so fast because it's all junk.
It's so many bugs, so many security issues, and blah blah blah. So yeah, it's amazing how quickly it moves, and how people at the same time they're concerned about their privacy, right? But at the same time, they also willing to give up and share all their data, all their information, just just to get the answers or just to ask one question or something like that. It's it's interesting.
Jonathan Schaeffer (42:08.953)
it's interesting, but I guess I'm old school, I value privacy, and I think we should all value privacy. But
I am concerned at the public perception of these programs. The media and the companies have sort of created this impression about how capable these programs are, how intelligent they are. And
It's really only in the last year that there's been a reckoning and people have understood that these programs have severe limitations. I remember in the first year after Chat GPT exploded on that magical day, November 30th, 2022, you know, companies said, wow, you know, we're gonna use these programs for our online help and we're gonna replace people with with programs that answer users' questions. And
many of those companies have gone back to using people because the programs, yeah, they got it right sometimes and they got it wrong lots of times, and those lots of times were pretty costly to the companies. People are discovering these programs have limitations. You know, a 95% accuracy rate sounds pretty darned impressive. Hey, if I was a student and I was scoring 95% on all my exams, I'd be really proud.
Vit Lyoshin (43:22.615)
Mm.
Jonathan Schaeffer (43:44.116)
But when it comes to AI, 95%, even 99% isn't good enough. We need to trust these programs that have that they do the right thing and they don't yet. And people say, come on, 99%, that's good enough. But think about it this way: you own a car, you have no problem driving it to and from work, going out on holidays, going to the theater. You might put
20,000, 30,000 miles on every year. You might have done, you know, a quarter of a million, a half million miles in your lifetime. You trust your car that it's safe and it's dependable. And part of the reason it is, is because there's agencies that oversee the cars. I just can't just build a car and be expect to sell it. That there's rules and regulations, it has to be inspected.
Vit Lyoshin (44:39.021)
Mm-hmm.
Jonathan Schaeffer (44:39.279)
If I design something and I said, hey, here's this car, and guess what? My brakes they only work 99% of the time. And I'll give you a really good price because they only work 99% of the time. You would never buy the car because 99% on brakes is a no-starter. You would never accept a car that only the brakes only work 99% of the time. Well, with AI, we're being conned to accepting AI.
Even though it lies. And I don't know what the error rate is, the fabrication rate, the hallucination rate is. It's high. It's you know, it's certainly much higher than one percent of the times. Let's let's be conservative and say it's five percent of the time. But again, do I want a product where five percent of the time I'm getting crap?
Fabricated references, made up stories, made up facts? the answer is no. And we we need a we need a mindset change. I would never accept 95 or 99% of my car. Why am I accepting it in my software products?
Vit Lyoshin (45:36.323)
Yeah.
Vit Lyoshin (45:54.692)
Yeah, and the problem is you don't know which five percent or one percent is is a wrong, right? If you knew that would be easy. If it would say it those five percent instead of lying or saying something, it would say I don't know. that would be acceptable. But yeah, that's a challenge. You don't know which five percent it is. So it's a hundred.
Jonathan Schaeffer (46:09.449)
Right.
Jonathan Schaeffer (46:16.381)
Well it's but again, you know yeah, but this is this is consistent with the fact that it does this coin flipping and so sometimes you get the right answer and sometimes you don't, and who knows when that is. But that's your responsibility. You're expected to look at the answer and do your own homework to see whether you trust the answer.
Vit Lyoshin (46:24.409)
Yeah.
Vit Lyoshin (46:32.975)
Yeah.
Vit Lyoshin (46:43.183)
Yeah.
Jonathan Schaeffer (46:43.835)
I would love to create a day where we have AI programs that we trust, that these programs work at such a high level that we have confidence that they're going to get it right. Everybody occasionally makes a mistake. It's human, but right now with these programs, the mistake level is just far too high, unacceptably high.
Vit Lyoshin (47:10.381)
Yeah. I wanted also to ask something else which related to privacy and and accuracy. Like we're being asked to provide as much context and details as possible for these AI programs to be more accurate for us. First of all, how true is that? And second of all, like what's the limit? How much you need to give it to get the actually the acceptable level of accuracy. I guess for many people it could be different, but
I mean I mean like how to balance it and can it even be possible to do that?
Jonathan Schaeffer (47:48.47)
yeah, th there are definitely AIs out there where they learn from your behavior, they better understand what you're what you do and what you like and what you don't like. And I guess humans are like that too. I mean, obviously, you know, if you're married, you learn what your wife likes and doesn't like, and she vice versa, and we we create this symbiotic relationship. And there are AIs out there that
try to do that. But, you know, again, it's not perfect. It's these computers aren't human. They don't understand empathy. they don't understand the psychology of humans well enough. And so it it's it it's it's hard for these programs to to learn. The other thing that is difficult is that
There's something amazing about the human brain, which I don't understand, is I can show you one example of something and you'll learn it immediately and never have to learn it again. If I if I was a kid and I went up to the oven and the the burner was on, and I touched the burner and I went, ouch, right? I would learn that that's hot. And if I touch it, I'm gonna get a sore.
Hand, finger, whatever, and I won't do it again because it was too painful. But I worked on poker programs for a number of years, look teaching AI how to play poker to be a strong poker player. And you just look at at these programs and how they learn, they need lots of statistics that they just
Vit Lyoshin (49:18.063)
Mm-hmm.
Jonathan Schaeffer (49:39.028)
They don't have the intuition, they don't have the understanding that humans do. And they don't learn from one data point. You need to give them lots of data points. And then eventually they start figuring out that, you know, maybe you're you like to bluff a lot or this or that or whatever. But they need lots of data before they they're confident in in making a decision. And
We're not like that. We can learn really, really quickly about other people and what their preferences are, how they operate. And these computers just need lots of data. And that's in part why these large language models want more data, right? The internet, everything that's on the internet isn't enough. They'd love to see every single file on everybody's desktop. They'd love to see all the data that's behind paywalls or in s secret locations or whatever, because the more data they get,
the the more they can make better coin flips. And if I can just I don't want it get philosophical or religious or whatever, but the more you study AI, the more you you just get amazed at at the human brain and what its capabilities are. it's unbelievable
That in such a small space right here, it A weighs what, about three pounds on average. and yet what we can do with speech and vision and hearing and reasoning and memory. you know, if I were to design the human brain, if I was the one who could design it, I would have had another slot in there so I could add a more memory, but given my limited memory, I think we all do.
Vit Lyoshin (51:09.612)
Yeah.
Vit Lyoshin (51:29.825)
Yeah.
Jonathan Schaeffer (51:30.688)
Pretty well. And it it's it's incredible. And what's even more amazing to me, and it's actually very satisfying, is when you see humans being able to do things very easily that computers find challenging or difficult. Anyway, I have enormous respect for for what we as a human race have have accomplished and will accomplish in the future with AI's help.
Vit Lyoshin (51:43.311)
Yeah.
Vit Lyoshin (51:50.445)
Yeah, it it
Vit Lyoshin (52:01.476)
Yeah. Yeah, I think it's a topic for another podcast because I immediately have a bunch of questions about that. How how to compare like computers and brains and how many years it takes to get to that level. So but just wrapping up the conversation here, from your opinion, what would be something that you personally like most people, or all everybody who use AI to adopt certain
behaviors or certain practices, interacting with the AI tools that you think will help them use it more, use it better, like don't give up too much privacy, but also get all this like you saying augmented intelligence from them.
Jonathan Schaeffer (52:50.973)
I think everybody needs a a basic AI education. And I don't mean that they have to go back to school and take a course, but just understanding that these programs can fabricate. And it's not malicious. It's not like they're trying to do you any harm. It's just fundamental i in how they are. I mean, I have some friends who lie a lot, and I know that. And
Vit Lyoshin (53:16.173)
Ha ha ha.
Jonathan Schaeffer (53:17.543)
I certainly take it a account when they make some claims that I clearly don't believe. But people just need to understand that
Vit Lyoshin (53:23.481)
Yeah.
Jonathan Schaeffer (53:30.388)
These programs are not intelligent. You shouldn't treat them as intelligent. You should treat them as tools that can help you. And the other thing that I think is really important is there's no free lunch. And what I mean by that is these companies are very happy to entice you with free things.
And it in many ways, maybe it all started with social media. Hey, get a free social media account and upload your photos and your videos and you can chat with your family and friends. But of course, it's zero cost to you. It's costing them real money because they need the infrastructure, et cetera. But they're getting data. And they use that data and they can mine that data.
And it's no different now with AI. They want to encourage you to put your data in the cloud. Right? We'll host your data, we'll do the backup of your data, but they get to see your data and they might train off your data and use your data. And you know, I've I've written books, for example, and the books they're not on the internet. I've never put my books on the internet. I guess maybe somebody has, but you know.
I've seen examples where AIs come back and they've plagiarized from my book. And I don't want to put my stuff up there and see it plagiarized. If I was a company, or here I'm a professor, I have my intellectual property, I have hundreds of notes of ideas I've had in the past. Companies have documents of their strategic plans, ideas, confidential relationships with customers. You put it up in the cloud or
Just send it that data to a chat program to analyze. There is a significant risk that that will be ingested in these programs and might show up somewhere else unexpected in ways that maybe don't reflect very well on you, or who knows what. And
Jonathan Schaeffer (55:46.012)
We need to be aware, just like using computers with networks, you know, we've had to learn a whole new n no nomenclature because there can be spam. How many people knew about spam until we got an email? And what about phishing attacks and denial of service attacks? I mean all these new terms that we've we've learned. and now with AI there's new terms and new risks that we need to be aware of. And I just want consumers
To be aware. I'm not saying don't use ChatGPT and Gemini and Claude and Lama and whatever. I'm just saying just be a knowledgeable user. Know what you're doing, that the answers may be wrong, how to recognize when they're wrong, know understand about data and how it might be used or abused. And
Just make the decisions that are best for you. I have lots of data that I don't care if it goes in the cloud because, you know, it's it's doesn't bother me. But I have lots of data that I never want to see in the cloud. I'm not gonna put my personal data in the cloud. I want it on my desktop and I wanna be able to use state-of-the-art tools to analyze it, for example.
Vit Lyoshin (57:03.327)
Mm-hmm. Yeah. Yeah, that's great. Yeah. All right. Thanks, Jonathan, for all the information and sharing your knowledge and wisdom with us. But before you go, I have three questions. My innovation QA. They're short. They're they're supposed to be short. So can you define innovation in a few words?
Jonathan Schaeffer (57:16.175)
Okay.
Jonathan Schaeffer (57:25.819)
innovation to me is tackling a real problem. Not not an abstract problem, but something that's real, that affects real people, that has practical value. So innovation is solving problems that affect real people. And so I would like to call my product kind a form of innovation.
Vit Lyoshin (57:51.83)
Mm-hmm. Okay, great. Second question is which innovation in the history of the world do you think changed the world the most?
Jonathan Schaeffer (58:03.123)
Well, history of the world's a long time. So if I just think about in my lifetime, I believe the most profound change, innovative change, came from the World Wide Web. Because it took this world and made it a lot smaller. We could communicate with people in rather cumbersome ways. You know, I could make a very expensive long-distance phone calls to my family.
Vit Lyoshin (58:08.847)
Okay.
Jonathan Schaeffer (58:32.979)
we had email, we had networking then, so I could send emails, but the world wide web really shrunk the world. I can I can literally go anywhere in the world and do almost anything from my desk, from my office, and it's amazing.
Vit Lyoshin (58:54.627)
Yeah. The last question is which one thing like a technology software or device that we use today we will be laughing at ten years from now?
Jonathan Schaeffer (59:08.467)
There's a lot of things, but
I would have to say it's passwords. you know, I have to right now come up with a password. I have sites which you have to have at least eight characters, and you have to have a digit in it, and a special character, and there's gotta be at least one capital and one small. And you're encouraged to have every site a different password. And I can't remember all of them. And
Vit Lyoshin (59:16.888)
Okay.
Jonathan Schaeffer (59:42.711)
on and on and on. Whereas retinal or fingerprint or some other form visual recognition, bang, I'm in, right? And so I have wasted a lot of time with passwords just because I don't remember them and I'm asked to change them often on many sites and it just find it frustrating. So I look forward to the day when I'm not typing in text tool passwords anymore.
Vit Lyoshin (59:44.591)
Mm-hmm.
Vit Lyoshin (01:00:12.193)
Yeah, I look forward to that too. That's a good one. All right. Thank you very much. And yeah, I hope we can talk more about other topics sometime in the future and stay in touch. And that's been great.
Jonathan Schaeffer (01:00:24.179)
Thank you very much. It it was a real pleasure talking to you. Appreciate the opportunity.
Vit Lyoshin (01:00:29.121)
Yeah, of course. Thank you very much. Bye bye.
Jonathan Schaeffer (01:00:31.968)
Goodbye.
YouTube
Spotify
Apple Podcasts
Amazon Music
iHeartRadio
Castbox
Overcast
RadioPublic