Karel Šafr: "I don't think there is real AI yet as we perceive it"
In today's podcast we introduce Karel Šafr - teacher, lecturer and researcher at Unicorn University. Karel has a long-standing interest in data processing and mathematical-statistical models. He teaches courses at the university that combine knowledge of mathematics, statistics and programming. Specifically, he teaches the courses Data Analysis and Visualization, Data Interpretation and Presentation, and the very popular Introduction to Machine Learning. As part of Unicorn University's professional courses, he is leading a new academy called Data Analyst in Python. In addition to teaching, he is working on Artificial Intelligence within the Unicorn AI research center. He also works at the University of Economics and the Czech Statistical Office.
Hi, in today's podcast we're going to talk with Karl Šafr, educator, lecturer and researcher at Unicorn University. Karel has a long-standing interest in data processing and mathematical-statistical models. He teaches courses at the university that combine knowledge of mathematics, statistics and programming. In addition to teaching, he is involved in artificial intelligence at the Unicorn Research Center and also works at the University of Economics and the Czech Statistical Office. Karl, thank you for coming.
Hi and thank you for having me.
I was wondering how you got to work at Unicorn University and what subjects do you teach there?
I got into this job through the current rector, Jan Čadil, who approached me if I would like to take up subjects that are mainly on the topic of data processing and analysis, and I am currently teaching these subjects. Specifically, it's Data Interpretation and Visualization, or I teach a very popular course called Introduction to Machine Learning, because statistical machine learning is a subject that is close to my heart.
What made you decide to pursue these particular subjects?
For me, it's a long-standing professional interest because I like to focus on topics like statistical machine learning or artificial intelligence as a whole. And overall, the topic of data mining, not only purely theoretically i.e. looking at models, but also practically seeing the use of these models, is very close to my heart. So for me it was an obvious path.
Apart from your university activities, you are involved in other projects, what are they?
I do quite a lot of research projects. I'm currently working on a couple of them. For example, they are projects on stock market derivatives prediction, where I work with students who then use this in their theses, and we are looking at using statistical machine learning to predict time series, or we are looking at data prediction and data desaggregation using statistical machine learning. I also work on the aforementioned artificial intelligence within the Unicorn AI Research Center (URC)[1] , where we also invent various new solutions directly for customers. We are involved in projects that are also handled by the company, so that it reaches the end customers as well.
[1] The Unicorn AI Research Center (URC) is a research center operating under the Unicorn University University, dedicated to research in the field of Artificial Intelligence (AI). For several years now, AI models have been experiencing a major boom, which we at the URC want to participate in. Behind the current wave of attention are large language models, in particular the popular ChatGPT tool, which offers many new possibilities but also reveals potential risks to keep in mind when developing AI.
How can these projects be combined with teaching at university?
That's what's interesting about it. As I'm involved in AI, it's a very popular topic these days, and it's growing dynamically, and we can see that most companies are using and want to use the full potential of AI, but the current job market is not yet saturated with people who can do it. As the school (Unicorn University) and the company (Unicorn) are connected, we are tapping into the potential that we have in the school and within the company. Researchers work with the company, which is very desirable then for getting work experience or for our students. There are still not enough people working on artificial intelligence and statistical machine learning and companies are commonly in demand from academia.
Which part fulfils you more? Is it teaching, consulting or working on projects?
I like a bit of everything. I'm the kind of person who wouldn't like to spend all week in one job, so I'm comfortable "bouncing around" doing something different every now and then. I like teaching a lot, and it may not seem like it, but it's quite draining, especially when you want to give it 100%. It's also demanding on the vocal cords, so I like it to be mixed like that. I can't say that I prefer just teaching or consulting or research.
I like the mix, which is very nice. At the same time, one thing is very important to mention here. Because artificial intelligence is a very dynamic field, what was here five years ago is often already outdated. You have to move very quickly with the times, and this combination allows me to do that. If I were only doing corporate projects, for example, I believe that I might get a bit marginal again in the other pages. Because not only is AI theory moving forward, but technology is moving forward quite dramatically, and one cannot close oneself off and focus on just one topic without somehow neglecting the others. Since this field is very dynamic, one can sleep for two or three years and then one is standing on the sidelines. So that's why I like this combination.
And do we even have a time when artificial intelligence is present?
That is a very good question. You could say that almost every day we see in the headlines the arrival of artificial intelligence, that artificial intelligence is going to take jobs and what all industries are threatened by artificial intelligence. I would venture a heretical idea that relatively few people have claimed, but it has been voiced. And that is that there is no artificial intelligence here yet. My view stems from what our expectations are. When you say AI, most people think of movie footage, some robot that has risen up and is looting the world. At best, it saves the world. Or they think of Halo 9000 from Space Odyssey. A robot that was part of a space ship that the crew talked to. They imagine some entity that's a separate entity that often has some superpowers from our perspective. And this view of ours has been articulated in movies and books, and I think it's very distorted. What artificial intelligence is is a little bit different. It's not some entity that decides on its own that, for example, my vacuum cleaner is going to get up one day and cut me in bed. I don't think so. The way I understand artificial intelligence, it's still statistical machine learning, and we have to remember the fundamental thing that it's still a mathematical model, which may be complex sometimes, but mostly it really just consists of matrix operations. Most of the time, derivatives come in, so some changes in functions, which is not relevant right now, but it's a mathematical model that's based on data. The kind of data that we put in the input, at best we can expect to get out of the output if the model is well done. I don't think there is any time here where we have artificial intelligence in the sense that we have learned to perceive it in the last 50-90 years from Isaac Asimov to other other books. And that they would be independent thinking units that would plan, make decisions, have some desires of their own. They don't. Most models are trained on data that have some function that they maximize, usually it's error minimization, and that error minimization is built on the data that we submit. Every artificial intelligence that we have here, whether it's the big language models that are very popular these days like ChatGPT or some of the others, they all have a specific purpose that a human has given them. And they work through data, a loss function and error minimization and they can't step outside of that framework that the researcher gives them. I think this is something that a lot of people don't realize and look at AI very broadly. The way we understand intelligence as a human being is the way we understand intelligence in a high-powered machine, and that's not the way it is. That time, I don't know if it's coming anytime in the future, is not here yet. I would venture to agree with the view that the time has not yet come when machines will rise.
How do you explain the recent popularity of AI? Where do you think it will be used the most?
It's a change similar in some areas to the tech revolution or the advent of photography during the 19th century, for example. It's a change we're experiencing in some way. And just as photography has influenced painting, AI will influence some other fields. It will certainly affect our everyday activities that AI handles. I certainly don't think that AI will completely replace all fields. A nice example is design. A lot of companies have recently jumped into generating product designs based on AI, and instead of having 50 designers where each one makes 10 designs and then the managers or some testers have 500 designs on their desk and decide which design is the best, typically for shoes, for example, AI can make the shoe designs on its own. And it will make 500, maybe 50,000 shoe designs, but then ultimately there still has to be a human who says this design here is nice and this design here is not nice. AI can't judge that a design is pretty. It takes the essence or patterns from nice designs and tries to extract those. Often AI lacks context and lacks some consideration of whether it is realistic. I'll give a nice example. Let's say that nice shoes for out in the woods would be shoes made of rough leather, proper laces, shoes with rough soles, and in this dataset we can have images of shoes for prom. And then you say to the AI create me nice shoes for prom, and it takes all the data and says, yeah here's a description of nice shoes, albeit for the woods, but nice shoes. The AI takes the sole from the canal and adds it to the prom shoes. So that can happen, too. Another nice example of what AI often does is it's not aware of context. It confuses things when it's analyzing photos. Our listeners can probably look up "Tiger vs. dog".
This is a rather nice photo of a labrador lying behind the fence bars, with the fence casting a shadow on him. The AI says there's a tiger because the colour the labrador has (stripes) is similar to a tiger. This is what the AI doesn't realize. It is unaware of the context. So I don't think she's going to replace or abolish entire fields, going back to designers, but she's going to have some implications into fields. It's going to have a consequence on streamlining work, it's going to have a consequence on the fact that it probably won't take as many designers to do a high volume of designs, but rather they'll be needed for specialized work like deciding that the product is even realistic to make. Eventually someone has to come in anyway to do the real technical re-drawing, to finish it and work out the details, because often AI gets lost in the details. So it's going to have an impact, but I don't think it's going to have such an impact that we're all going to lose our jobs and sit at home. After all, human labour has added value that AI can't capture.
Thank you for that answer.
Where do you see the threats and risks of AI?
We encounter threats and risks almost daily in newspaper articles. Of course, there are quite a few threats and risks. As big language models emerged, so did the use in viruses. What hackers were doing was uploading the source code of a virus into a language model and having the language model rework that source code. Why? Because, for example, antiviruses work on the basis that they analyse the structure of the source code, and when they find something like that, they declare it a virus, but as the source code was reworked and had a different structure, it didn't look like a virus right off the bat. Another very negative use, currently being discussed in academia, is the use of language models to create theses. For this reason, some universities are gradually abandoning the concept of theses, so far only in undergraduate theses, and replacing them with student projects, for example. Of course, these are impacts that are negative, they are a threat, and society has to react to them somehow. The biggest problem I see is whether we can react to it. Whether the legislation, which is often a long-distance thing, can respond adequately. We observe two situations around us. Some countries are overreacting, which can sometimes be a negative thing, that the situation is overregulated. Or they are not reacting at all and are opening up the Wild West. Which, again, given that AI systems should receive our personal data, let's say not only in Europe but abroad, that could also have a very negative impact. So there are threats mainly in terms of society being able to react and get used to it. Take, for example, technologies that can mimic the human voice, so that someone could make a five-minute voice recording of you and, based on the recording, be able to read any text in your voice. That's a very negative thing. You can see that this abuse is slowly happening and you can see it especially abroad. Or the abuse of fake videos i.e. Deepfake[2] - that is a big topic. The overall post-reality is that we are in a time when we start having problems to distinguish what is reality and what is not. I think all of these problems can be overcome, we're just going to have to go through a time where we're going to have to change a lot, or rather refine our perception of the world in the sense of being aware of all of this and being able to work with it.
[2] Deepfake video is a technology that uses artificial intelligence, specifically deep learning, to create realistic but fake videos. These videos typically manipulate images and sound to show people doing or saying things they never actually did or said.
Do you think AI will take jobs away from programmers?
Here I would go back to the example where the photo came up. All the newspaper articles at the time claimed that the study of painting and art in general was a field that would disappear and painters would lose their jobs. This is now happening again with AI and similar headlines are appearing from time to time. In the end, that turned out not to be the case and the 20th century was the biggest boom painting art has seen. The pressure that photography created on painting was such that new painting styles emerged and it could be said that 20th century painting is one of the most interesting periods. Until then it was more realism, but new artistic styles emerged that focused on design or abstraction. Photography influenced painting and forced it to evolve somewhere else. I think this is what programming is waiting for. Not only programmers, but also testers and in general all IT fields are going to move in some direction. I think it's going to be a tool that we as programmers will use, but I don't think it's going to take away our jobs. It's more like it's going to be some help, some colleague that we can ask about how to create a feature or how to implement something or help us design some tests. It will be a colleague who we will not be able to trust 100% because he sometimes makes mistakes and sometimes lies. Language models in particular tend to lie, and lie very convincingly. He sometimes mixes the fifth with the ninth, so I don't think AI will take away programmers' jobs. I think it's going to make the fields more efficient and take them further. I think it's more of a tool that programmers will have to learn to work with. I think human labor is still very much valued, and that's what happened with painting. Even though we have photography, digital techniques, and we can ask AI to create a painting, it still turns out that classical painting and painting are also very valued fields. We can go to the National Gallery, DOX or anywhere else to see exhibitions of contemporary artists and we will find that firstly it is still very popular and secondly these artistic things are very much valued.
I'm wondering, what's the deal with AI in school? How are students using AI in their studies and what do you think about the issue of writing theses?
That's a topic for a whole conversation. Nowadays, a lot depends on the individual school's approach. There are schools that choose a completely repressive approach, i.e. they block access to these technologies not only in the Czech Republic but also abroad - England, Italy and other big countries. The second is to somehow allow it or let it into the classroom, either in a limited mode or in some kind of maximum mode, as much as possible. Personally, I don't think those repressive approaches have a chance of lasting in the long run, and going forward we need to realize that not only students but all of us will need AI.
On the contrary, I think you have to learn to work with it. I'm trying to make it somehow accessible, even in student work. It's important to realize that AI often doesn't do exactly what humans ask it to do. Now we're talking about prompt engineering, which is the name for being able to give commands so that the AI does what we want it to do. I think this is a trait that we as humans will have to learn i.e. how to interact with AI, how to read the results and also how much to trust it. I think there is a point in teaching, but there is a need for education and working with facts. Just as we have education for students, especially in high school and elementary school, about media skills so that students are able to work critically with information, I think it's important to educate about AI. That students learn to use AI tools sparingly and in the context of the purpose, because these tools are not made to write entire papers or create complex projects, but they're often trained on data sets that are very short, so it's more of a question and answer kind of thing. I think, especially language models end up in some better encyclopedias, but which we can't trust 100%.
How's Unicorn University? In what courses can students encounter AI?
We allow students to use AI in the writing of some projects. This is individual, and always depends on the instructors and the context of the assignment. We can't say that it's desirable everywhere. For example, I use it and allow it within some final assignments or projects. In the context of tests, that's where it shouldn't be used, so that it doesn't tell you the straight answer. Because it's one thing to use it to learn something and it's another thing to use it to cheat. So if I have to demonstrate some knowledge of my own, it's usually not desirable to use AI. Conversely, if I have to demonstrate some ability to create something and it's a project and the teacher in question considers it a normal tool, then we use it. It's very individual and you can't say for all schools these days that it's allowed and used everywhere. Let's face it, the purpose of a test is to test knowledge and knowledge should not be based on some probability model, but the person should have it, so for example on tests using AI doesn't make sense.
You mentioned that there are theses on AI. Can you give us an example?
I have several students who are doing their theses on AI. I have a student who is looking at time series predictions in stock market operations. Another nice example that I think is interesting is when our student who is analyzing a court judgment using language models creates their own or modifies existing language models that can process a court judgment. So he would pull in information and background on the court proceedings and based on that, the model would recommend a court action or predict how the court proceedings would turn out. Next, quite a lot of students are doing a popular thing, and that's creating opponents for computer games, so some sort of engine in the background of a computer game that allows you to play against the computer.
What AI-related projects have Unicorn University students worked on?
These are projects where we create an opponent for a computer game in workshops. So I'd like to play a demonstration of the cars here. Now we can see the cars driving around the computer. The red blocks are the obstacles that they hit. The first time they went out, they all crashed right away. And we can see that the longer they drive, the better the cars are able to drive. Now they've been driving almost all the time. The goal of the cars is to drive through as many blue gates as possible without hitting any obstacles. We may notice that after a while the cars will be able to drive through or they will drive almost endlessly around the game board and have no problem crashing into obstacles. The interesting thing here is that there is an artificial intelligence model or rather neural networks in the background that helps the cars make decisions. The very first generation of cars that drove there had completely random coefficients and now there's a car that systematically drives through the gate and doesn't crash. Here we use what we call an evolutionary algorithm and other tools to create the behaviour that is desirable. Here, in this case, it's a car that will drive and not crash, and it will drive through the maximum number of blue gates. It's a project that students often work on in a course like Programming for Artificial Intelligence, an introduction to machine learning. Then I'd like to ask you to start the second project, which is the sandbox. Here we see live training of artificial intelligence. Artificial intelligence learning to play ping-pong. The goal is to create an AI-based game engine that will be able to play against a human in a game of Mouseketeers. Here, I teach students the ideas and principles of creating such game engines while working with them on data, actual implementation, and other specifics. Game engines often need to be trained on several difficulties, i.e. not to defeat the opponent right away or not to make it too easy. Which is something we encounter relatively daily and don't even know it. Especially people who play chess on their phone or other games, there are usually some opponents and in a lot of cases they are trained by some AI tool. It's very common in chess, but it's not just chess, it's also like racing simulators or strategy games where there's a choice to be made. It's a very popular topic among students. I have a student right now who is looking at the potential use of artificial intelligence in the card game of blackjack.
What is the future of AI at Unicorn University?
We are currently developing new courses to be taught. It's game engines that are very popular, so it will be the use of AI in game engines. We are also preparing a new undergraduate program on AI, where students will learn about machine learning and AI. This will allow them to understand the basic principles of how data is processed and how AI is created from that.
Other courses you teach are Data Interpretation and Presentation and Data Analysis and Visualization. Do you think these subjects are an essential component for education in IT and economics?
I think they definitely are. Understanding data is very important. It's not just about looking at data through patterns or tables, printing out the data or looking at the patterns that come out of the data when we calculate the average. I think the essential thing is visualizing the data, i.e. being able to make some kind of graph on the data and somehow prepare it, transform it, because the human brain doesn't completely understand data through formulas, but usually understands data through an image or some kind of visual perception. A nice example of that is constellations, if we look at the sky. People over the centuries have learned to perceive the stars through constellations and have learned to navigate the stars, for example, they knew where Orion was, where various other kinds of stars like Polaris were, and they knew and how to navigate using constellations. However, if we gave people a piece of paper with the coordinates of the stars in the sky written on it, they would never form the reasoning that we have a cluster of stars, that there are the Pleiades, etc. And that works in data analysis as well. Whether it's some financial data, IT data, stock market time series, bitcoin trends, or maybe GDP trends, I think it's terribly crucial to be able to understand the data and be able to work with it when I'm trying to make a model that will predict or trying to sort and classify the data. The role with data can be all sorts of things, but it's very important to be able to plot the data, process it properly and pick the right charts. There are an incredible number of charts these days, not just the basic dot or line chart. The potential that is behind graphs is huge and we are very much involved in that at the school for the reason that students are able to visualise and process data very effectively and understand what the limits of visualisations are. Sometimes what happens is that people make the wrong graphs to be able to understand what the potential of the data is. In data interpretation and presentation, we cover things like dashboards, graphing rules for creating dashboards, and more.
First of all, you teach in master's degree programmes. Is this study more demanding than the Bachelor's degree? Can you tell us the reason why students should pursue further studies?
Masters studies continue in some way and it is a natural evolution. What a person learns in an undergraduate program, he or she builds on those subjects further and further. At the same time, I don't think it's something unrealistic to say, "Now I've done my BA and I'm scared to go to graduate school." The master's degree that we offer at the school is structured so that students can build seamlessly on the knowledge gained from the bachelor's degree. To allow them to further their education and continue what they have learned to gain extra knowledge. In this day and age AI and data processing in general is a very important and big area, so I think there's a huge added value for the student, especially in being able to work with data and I think the job market can appreciate that. Which is evident in data analysts, who are in high demand. A data analyst who can code is a very desirable profession and very well regarded.
Among other things, you are a lecturer of a professional course from Unicorn University - Data Analysis in Python. Can you introduce this course? Who is it suitable for?
The course is especially suitable for people who already know the basics of Python. We offer two courses and mine is already for more advanced. If someone doesn't know the basics of Python, they can start with the course offered by my colleague. My course is primarily suitable for those who already know the basics. In today's world, Python and R are some of the most widely used data processing languages. In this course, we cover Python and in the future we will offer a course on R - R data processing languages as well. The Data Analysis in Python course is highly desirable due to the high popularity of the language itself. The course is all about explaining the basic principles of how data is processed, so it's not just about how to retrieve the data, but also cleaning the data, i.e. removing unwanted phenomena and outliers. We also cover the basic libraries that are most in demand in the job market today, such as the NumPy library, or Pandas i.e. a library for spreadsheet data grabbing and processing. We are also dedicated to libraries for analyzing our own data using mathematical and statistical models, so from basic regression models to technologies such as Random Forest or Adaboost, or we are dedicated to artificial intelligence, where we work with tools that are nowadays, both in the private and academic sphere, one of the most used tools for creating artificial intelligence models. We explain how to work with these tools and what are the basic principles of libraries to create a basic or even more advanced AI mo
Our conversation is coming to a close. I wonder what fields have the best prospects in the future and what would you study if you could choose again?
That's a good question. I probably wouldn't choose a field that is professionally focused on just one particular detail, but I would definitely want one that combines knowledge of multiple fields of programming, statistics and data analysis, which is the field of artificial intelligence, for example, because it's a combination of several fields: IT, statistics and mathematics, so I think I would choose something like that. And I find it the most interesting and the most promising, because if you focus only on one particular detail, for example, one particular language, and limit your perception of the world to one part, then after a while it may happen that it is no longer in demand on the job market. So I think the field of AI would attract me the most.
Our guest today was Karel Šafr, AI engineer, educator and university lecturer. Karel, thank you very much.
Thank you for having me and thank you for the interview.