{"id":311554,"date":"2025-11-22T16:03:19","date_gmt":"2025-11-22T16:03:19","guid":{"rendered":"https:\/\/peraltafinancing.com\/uncategorized\/the-llmops-shift-with-abi-aryan-oreilly\/"},"modified":"2025-11-22T16:03:19","modified_gmt":"2025-11-22T16:03:19","slug":"the-llmops-shift-with-abi-aryan-oreilly","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=311554","title":{"rendered":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"postContent-content\">\n<div class=\"podcast_player\">\n<div id=\"2962195828\" class=\"castos-player dark-mode \" tabindex=\"0\" data-episode=\"17736\" data-player_id=\"2962195828\">\n<div class=\"player\">\n<div class=\"player__main\">\n<div class=\"player__artwork player__artwork-17736\">\n<img decoding=\"async\" src=\"https:\/\/www.oreilly.com\/radar\/wp-content\/uploads\/sites\/3\/2024\/01\/Podcast_Cover_GenAI_in_the_Real_World-160x160.png\" alt=\"Generative AI in the Real World\" title=\"Generative AI in the Real World\"\/><\/div>\n<div class=\"player__body\">\n<div class=\"currently-playing\">\n<p>\nGenerative AI in the Real World<\/p>\n<p>Generative AI in the Real World: The LLMOps Shift with Abi Aryan<\/p>\n<\/div>\n<div class=\"play-progress\">\n<div class=\"play-pause-controls\">\n<button title=\"Play\" aria-label=\"Play Episode\" aria-pressed=\"false\" class=\"play-btn\"><br \/>\n<span class=\"screen-reader-text\">Play Episode<\/span><br \/>\n<\/button><br \/>\n<button title=\"Pause\" aria-label=\"Pause Episode\" aria-pressed=\"false\" class=\"pause-btn hide\"><br \/>\n<span class=\"screen-reader-text\">Pause Episode<\/span><br \/>\n<\/button><br \/>\n<img decoding=\"async\" src=\"https:\/\/www.oreilly.com\/radar\/wp-content\/plugins\/seriously-simple-podcasting\/assets\/css\/images\/player\/images\/icon-loader.svg\" alt=\"Loading\" class=\"ssp-loader hide\"\/><\/div>\n<div>\n<audio preload=\"none\" class=\"clip clip-17736\"><source src=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3\"\/><\/audio><\/p>\n<div class=\"ssp-playback playback\">\n<p>\n<button class=\"player-btn player-btn__volume\" title=\"Mute\/Unmute\"><br \/>\n<span class=\"screen-reader-text\">Mute\/Unmute Episode<\/span><br \/>\n<\/button><br \/>\n<button data-skip=\"-10\" class=\"player-btn player-btn__rwd\" title=\"Rewind 10 seconds\"><br \/>\n<span class=\"screen-reader-text\">Rewind 10 Seconds<\/span><br \/>\n<\/button><br \/>\n<button data-speed=\"1\" class=\"player-btn player-btn__speed\" title=\"Playback Speed\" aria-label=\"Playback Speed\">1x<\/button><br \/>\n<button data-skip=\"30\" class=\"player-btn player-btn__fwd\" title=\"Fast Forward 30 seconds\"><br \/>\n<span class=\"screen-reader-text\">Fast Forward 30 seconds<\/span><br \/>\n<\/button><\/p>\n<p>\n<time class=\"ssp-timer\">00:00<\/time><br \/>\n<span>\/<\/span><br \/>\n<time class=\"ssp-duration\" datetime=\"PT0H0M0S\">32m 16s<\/time><\/p>\n<\/div>\n<\/div>\n<\/div>\n<nav class=\"player-panels-nav\">\n<button class=\"subscribe-btn\" id=\"subscribe-btn-17736\" title=\"Subscribe\">Subscribe<\/button><br \/>\n<button class=\"share-btn\" id=\"share-btn-17736\" title=\"Share\">Share<\/button><\/nav>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p>MLOps is dead. Well, not really, but for many the job is evolving into LLMOps. In this episode, Abide AI founder and <em>LLMOps<\/em> author Abi Aryan joins Ben to discuss what LLMOps is and why it\u2019s needed, particularly for agentic AI systems. Listen in to hear why LLMOps requires a new way of thinking about observability, why we should spend more time understanding human workflows before mimicking them with agents, how to do FinOps in the age of generative AI, and more.<\/p>\n<p>About the <em>Generative AI in the Real World<\/em> podcast: In 2023, ChatGPT put AI on everyone\u2019s agenda. In 2025, the challenge will be turning those agendas into reality. In <em>Generative AI in the Real World<\/em>, Ben Lorica interviews leaders who are building with AI. Learn from their experience to help put AI to work in your enterprise.<\/p>\n<p>Check out <a href=\"https:\/\/learning.oreilly.com\/playlists\/42123a72-1108-40f1-91c0-adbfb9f4983b\/?_gl=1*z3nry*_ga*MTcyMjYzMjI1NS4xNzYyOTU5MzYw*_ga_092EL089CH*czE3NjI5NjU4ODkkbzIkZzEkdDE3NjI5NjU5NzUkajM0JGwwJGgw\" target=\"_blank\" rel=\"noreferrer noopener\">other episodes<\/a> of this podcast on the O\u2019Reilly learning platform.<\/p>\n<h2 class=\"wp-block-heading\">Transcript<\/h2>\n<p><em>This transcript was created with the help of AI and has been lightly edited for clarity.<\/em><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=0\" target=\"_blank\" rel=\"noreferrer noopener\">00.00<\/a>: <strong>All right, so today we have Abi Aryan. She is the author of the <\/strong><a href=\"https:\/\/learning.oreilly.com\/library\/view\/llmops\/9781098154196\/\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>O\u2019Reilly book on LLMOps<\/strong><\/a><strong> as well as the founder of Abide AI. So, Abi, welcome to the podcast.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=19\" target=\"_blank\" rel=\"noreferrer noopener\">00.19<\/a>: Thank you so much, Ben.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=21\" target=\"_blank\" rel=\"noreferrer noopener\">00.21<\/a>: <strong>All right. Let\u2019s start with the book, which I confess, I just cracked open: <em>LLMOps<\/em>. People probably listening to this have heard of MLOps. So at a high level, the models have changed: They\u2019re bigger, they\u2019re generative, and so on and so forth. So since you\u2019ve written this book, have you seen a wider acceptance of the need for LLMOps?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=51\" target=\"_blank\" rel=\"noreferrer noopener\">00.51<\/a>: I think more recently there are more infrastructure companies. So there was a conference happening recently, and there was this sort of perception or messaging across the conference, which was \u201cMLOps is dead.\u201d Although I don\u2019t agree with that.\u00a0<\/p>\n<p>There\u2019s a big difference that companies have started to pick up on more recently, as the infrastructure around the space has sort of started to improve. They\u2019re starting to realize how different the pipelines were that people managed and grew, especially for the older companies like Snorkel that were in this space for years and years before large language models came in. The way they were handling data pipelines\u2014and even the observability platforms that we\u2019re seeing today\u2014have changed tremendously.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=100\" target=\"_blank\" rel=\"noreferrer noopener\">01.40<\/a>: <strong>What about, Abi, the general.\u00a0.\u00a0.? We don\u2019t have to go into specific tools, but we can if you want. But, you know, if you look at the old MLOps person and then fast-forward, this person is now an LLMOps person. So on a day-to-day basis [has] their suite of tools changed?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=121\" target=\"_blank\" rel=\"noreferrer noopener\">02.01<\/a>: Massively. I think for an MLOps person, the focus was very much around \u201cThis is my model. How do I containerize my model, and how do I put it in production?\u201d That was the entire problem and, you know, most of the work was around \u201cCan I containerize it? What are the best practices around how I arrange my repository? Are we using templates?\u201d\u00a0<\/p>\n<p>Drawbacks happened, but not as much because most of the time the stuff was tested and there was not too much indeterministic behavior within the models itself. Now that has changed.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=158\" target=\"_blank\" rel=\"noreferrer noopener\">02.38<\/a>: [For] most of the LLMOps engineers, the biggest job right now is doing FinOps really, which is controlling the cost because the models are massive. The second thing, which has been a big difference, is we have shifted from \u201cHow can we build systems?\u201d to \u201cHow can we build systems that can perform, and not just perform technically but perform behaviorally as well?\u201d: \u201cWhat is the cost of the model? But also what is the latency? And see what\u2019s the throughput looking like? How are we managing the memory across different tasks?\u201d\u00a0<\/p>\n<p>The problem has really shifted when we talk about it.\u00a0.\u00a0. So a lot of focus for MLOps was \u201cLet\u2019s create fantastic dashboards that can do everything.\u201d Right now it\u2019s no matter which dashboard you create, the monitoring is really very dynamic.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=212\" target=\"_blank\" rel=\"noreferrer noopener\">03.32<\/a>: <strong>Yeah, yeah. As you were talking there, you know, I started thinking, yeah, of course, obviously now the inference is essentially a distributed computing problem, right? So that was not the case before. Now you have different phases even of the computation during inference, so you have the prefill phase and the decode phase. And then you might need different setups for those.\u00a0<\/strong><\/p>\n<p><strong>So anecdotally, Abi, did the people who were MLOps people successfully migrate themselves? Were they able to upskill themselves to become LLMOps engineers?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=254\" target=\"_blank\" rel=\"noreferrer noopener\">04.14<\/a>: I know a couple of friends who were MLOps engineers. They were teaching MLOps as well\u2014Databricks folks, MVPs. And they were now transitioning to LLMOps.<\/p>\n<p>But the way they started is they started focusing very much on, \u201cCan you do evals for these models? They weren\u2019t really dealing with the infrastructure side of it yet. And that was their slow transition. And right now they\u2019re very much at that point where they\u2019re thinking, \u201cOK, can we make it easy to just catch these problems within the model\u2014inferencing itself?\u201d<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=289\" target=\"_blank\" rel=\"noreferrer noopener\">04.49<\/a>: A lot of other problems still stay unsolved. Then the other side, which was like a lot of software engineers who entered the field and became AI engineers, they have a much easier transition because software.\u00a0.\u00a0. The way I look at large language models is not just as another machine learning model but literally like software 3.0 in that way, which is it\u2019s an end-to-end system that will run independently.<\/p>\n<p>Now, the model isn\u2019t just something you plug in. The model is the product tree. So for those people, most software is built around these ideas, which is, you know, we need a strong cohesion. We need low coupling. We need to think about \u201cHow are we doing microservices, how the communication happens between different tools that we\u2019re using, how are we calling up our endpoints, how are we securing our endpoints?\u201d<\/p>\n<p>Those questions come easier. So the system design side of things comes easier to people who work in traditional software engineering. So the transition has been a little bit easier for them as compared to people who were traditionally like MLOps engineers.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=359\" target=\"_blank\" rel=\"noreferrer noopener\">05.59<\/a>: <strong>And hopefully your book will help some of these MLOps people upskill themselves into this new world.<\/strong><\/p>\n<p><strong>Let\u2019s pivot quickly to agents. Obviously it\u2019s a buzzword. Just like anything in the space, it means different things to different teams. So how do you distinguish agentic systems yourself?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=384\" target=\"_blank\" rel=\"noreferrer noopener\">06.24<\/a>: There are two words in the space. One is agents; one is agent workflows. Basically agents are the components really. Or you can call them the model itself, but they\u2019re trying to figure out what you meant, even if you forgot to tell them. That\u2019s the core work of an agent. And the work of a workflow or the workflow of an agentic system, if you want to call it, is to tell these agents what to actually do. So one is responsible for execution; the other is responsible for the planning side of things.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=422\" target=\"_blank\" rel=\"noreferrer noopener\">07.02<\/a>: <strong>I think sometimes when tech journalists write about these things, the general public gets the notion that there\u2019s this monolithic model that does everything. But the reality is, most teams are moving away from that design as you, as you describe.<\/strong><\/p>\n<p><strong>So they have an agent that acts as an orchestrator or planner and then parcels out the different steps or tasks needed, and then maybe reassembles in the end, right?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=462\" target=\"_blank\" rel=\"noreferrer noopener\">07.42<\/a>: Coming back to your point, it\u2019s now less of a problem of machine learning. It\u2019s, again, more like a distributed systems problem because we have multiple agents. Some of these agents will have more load\u2014they will be the frontend agents, which are communicating to a lot of people. Obviously, on the GPUs, these need more distribution.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=482\" target=\"_blank\" rel=\"noreferrer noopener\">08.02<\/a>: And when it comes to the other agents that may not be used as much, they can be provisioned based on \u201cThis is the need, and this is the availability that we have.\u201d So all of that provisioning again is a problem. The communication is a problem. Setting up tests across different tasks itself within an entire workflow, now that becomes a problem, which is where a lot of people are trying to implement context engineering. But it\u2019s a very complicated problem to solve.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=511\" target=\"_blank\" rel=\"noreferrer noopener\">08.31<\/a>: <strong>And then, Abi, there\u2019s also the problem of compounding reliability. Let\u2019s say, for example, you have an agentic workflow where one agent passes off to another agent and yet to another third agent. Each agent may have a certain amount of reliability, but it compounds over time. So it compounds across this pipeline, which makes it more challenging.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=542\" target=\"_blank\" rel=\"noreferrer noopener\">09.02<\/a>: And that\u2019s where there\u2019s a lot of research work going on in the space. It\u2019s an idea that I\u2019ve talked about in the book as well. At that point when I was writing the book, especially chapter four, in which a lot of these were described, most of the companies right now are [using] monolithic architecture, but it\u2019s not going to be able to sustain as we go towards application.<\/p>\n<p>We have to go towards a microservices architecture. And the moment we go towards microservices architecture, there are a lot of problems. One will be the hardware problem. The other is consensus building, which is.\u00a0.\u00a0.\u00a0<\/p>\n<p>Let\u2019s say you have three different agents spread across three different nodes, which would be running very differently. Let\u2019s say one is running on an edge one hundred; one is running on something else. How can we achieve consensus if even one of the nodes ends up winning? So that\u2019s open research work [where] people are trying to figure out, \u201cCan we achieve consensus in agents based on whatever answer the majority is giving, or how do we really think about it?\u201d It should be set up at a threshold at which, if it\u2019s beyond this threshold, then you know, this perfectly works.<\/p>\n<p>One of the frameworks that is trying to work in this space is called <a href=\"https:\/\/github.com\/massgen\/MassGen\" target=\"_blank\" rel=\"noreferrer noopener\">MassGen<\/a>\u2014they\u2019re working on the research side of solving this problem itself in terms of the tool itself.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=631\" target=\"_blank\" rel=\"noreferrer noopener\">10.31<\/a>: <strong>By the way, even back in the microservices days in software architecture, obviously people went overboard too. So I think that, as with any of these new things, there\u2019s a bit of trial and error that you have to go through. And the better you can test your systems and have a setup where you can reproduce and try different things, the better off you are, because many times your first stab at designing your system may not be the right one. Right?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=668\" target=\"_blank\" rel=\"noreferrer noopener\">11.08<\/a>: Yeah. And I\u2019ll give you two examples of this. So AI companies tried to use a lot of agentic frameworks. You know people have used Crew; people have used n8n, they\u2019ve used.\u00a0.\u00a0.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=685\" target=\"_blank\" rel=\"noreferrer noopener\">11.25<\/a>: <strong>Oh, I hate those! Not I hate.\u00a0.\u00a0. Sorry. Sorry, my friends and crew.<\/strong>\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=690\" target=\"_blank\" rel=\"noreferrer noopener\">11.30<\/a>: And 90% of the people working in this space seriously have already made that transition, which is \u201cWe are going to write it ourselves.\u00a0<\/p>\n<p>The same happened for evaluation: There were a lot of evaluation tools out there. What they were doing on the surface is literally just tracing, and tracing wasn\u2019t really solving the problem\u2014it was just a beautiful dashboard that doesn\u2019t really serve much purpose. Maybe for the business teams. But at least for the ML engineers who are supposed to debug these problems and, you know, optimize these systems, essentially, it was not giving much other than \u201cWhat is the error response that we\u2019re getting to everything?\u201d<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=728\" target=\"_blank\" rel=\"noreferrer noopener\">12.08<\/a>: So again, for that one as well, most of the companies have developed their own evaluation frameworks in-house, as of now. The people who are just starting out, obviously they\u2019ve done. But most of the companies that started working with large language models in 2023, they\u2019ve tried every tool out there in 2023, 2024. And right now more and more people are staying away from the frameworks and launching and everything.<\/p>\n<p>People have understood that most of the frameworks in this space are not superreliable.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=761\" target=\"_blank\" rel=\"noreferrer noopener\">12.41<\/a>: <strong>And [are] also, honestly, a bit bloated. They come with too many things that you don\u2019t need in many ways.\u00a0.\u00a0.<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=774\" target=\"_blank\" rel=\"noreferrer noopener\">12:54<\/a>: Security loopholes as well. So for example, like I reported one of the security loopholes with LangChain as well, with LangSmith back in 2024. So those things obviously get reported by people [and] get worked on, but the companies aren\u2019t really proactively working on closing those security loopholes.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=795\" target=\"_blank\" rel=\"noreferrer noopener\">13.15<\/a>: <strong>Two open source projects that I like that are not specifically agentic are DSPy and BAML. Wanted to give them a shout out. So this point I\u2019m about to make, there\u2019s no easy, clear-cut answer. But one thing I noticed, Abi, is that people will do the following, right? I\u2019m going to take something we do, and I\u2019m going to build agents to do the same thing. But the way we do things is I have a\u2014I\u2019m just making this up\u2014I have a project manager and then I have a designer, I have role B, role C, and then there\u2019s certain emails being exchanged.<\/strong><\/p>\n<p><strong>So then the first step is \u201cLet\u2019s replicate not just the roles but kind of the exchange and communication.\u201d And sometimes that actually increases the complexity of the design of your system because maybe you don\u2019t need to do it the way the humans do it. Right? Maybe if you go to automation and agents, you don\u2019t have to over-anthropomorphize your workflow. Right. So what do you think about this observation?<\/strong>\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=871\" target=\"_blank\" rel=\"noreferrer noopener\">14.31<\/a>: A very interesting analogy I\u2019ll give you is people are trying to replicate intelligence without understanding what intelligence is. The same for consciousness. Everybody wants to replicate and create consciousness without understanding consciousness. So the same is happening with this as well, which is we are trying to replicate a human workflow without really understanding how humans work.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=895\" target=\"_blank\" rel=\"noreferrer noopener\">14.55<\/a>: <strong>And sometimes humans may not be the most efficient thing. Like they exchange five emails to arrive at something.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=904\" target=\"_blank\" rel=\"noreferrer noopener\">15.04<\/a>: And humans are never context defined. And in a very limiting sense. Even if somebody\u2019s job is to do editing, they\u2019re not just doing editing. They are looking at the flow. They are looking for a lot of things which you can\u2019t really define. Obviously you can over a period of time, but it needs a lot of observation to understand. And that skill also depends on who the person is. Different people have different skills as well. Most of the agentic systems right now, they\u2019re just glorified Zapier IFTTT routines. That\u2019s the way I look at them right now. The if recipes: If this, then that.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=948\" target=\"_blank\" rel=\"noreferrer noopener\">15.48<\/a>: <strong>Yeah, yeah. Robotic process automation I guess is what people call it. The other thing that people I don\u2019t think understand just reading the popular tech press is that agents have levels of autonomy, right? Most teams don\u2019t actually build an agent and unleash it full autonomous from day one.<\/strong><\/p>\n<p><strong>I mean, I guess the analogy would be in self-driving cars: They have different levels of automation. Most enterprise AI teams realize that with agents, you have to kind of treat them that way too, depending on the complexity and the importance of the workflow.\u00a0<\/strong><\/p>\n<p><strong>So you go first very much a human is involved and then less and less human over time as you develop confidence in the agent.<\/strong><\/p>\n<p><strong>But I think it\u2019s not good practice to just kind of let an agent run wild. Especially right now.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1016\" target=\"_blank\" rel=\"noreferrer noopener\">16.56<\/a>: It\u2019s not, because who\u2019s the person answering if the agent goes wrong? And that\u2019s a question that has come up often. So this is the work that we\u2019re doing at Abide really, which is trying to create a decision layer on top of the knowledge retrieval layer.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1027\" target=\"_blank\" rel=\"noreferrer noopener\">17.07<\/a>: Most of the agents which are built using just large language models.\u00a0.\u00a0. LLMs\u2014I think people need to understand this part\u2014are fantastic at knowledge retrieval, but they do not know how to make decisions. If you think agents are independent decision makers and they can figure things out, no, they cannot figure things out. They can look at the database and try to do something.<\/p>\n<p>Now, what they do may or may not be what you like, no matter how many rules you define across that. So what we really need to develop is some sort of symbolic language around how these agents are working, which is more like trying to give them a model of the world around \u201cWhat is the cause and effect, with all of these decisions that you\u2019re making? How do we prioritize one decision where the.\u00a0.\u00a0.? What was the reasoning behind that so that entire decision making reasoning here has been the missing part?\u201d<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1082\" target=\"_blank\" rel=\"noreferrer noopener\">18.02<\/a>: <strong>You brought up the topic of observability. There\u2019s two schools of thought here as far as agentic observability. The first one is we don\u2019t need new tools. We have the tools. We just have to apply [them] to agents. And then the second, of course, is this is a new situation. So now we need to be able to do more.\u00a0.\u00a0. The observability tools have to be more capable because we\u2019re dealing with nondeterministic systems.<\/strong><\/p>\n<p><strong>And so maybe we need to capture more information along the way. Chains of decision, reasoning, traceability, and so on and so forth. Where do you fall in this kind of spectrum of we don\u2019t need new tools or we need new tools?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1128\" target=\"_blank\" rel=\"noreferrer noopener\">18.48<\/a>: We don\u2019t need new tools, but we certainly need new frameworks, and especially a new way of thinking. Observability in the MLOps world\u2014fantastic; it was just about tools. Now, people have to stop thinking about observability as just visibility into the system and start thinking of it as an anomaly detection problem. And that was something I\u2019d written in the book as well. Now it\u2019s no longer about \u201cCan I see what my token length is?\u201d No, that\u2019s not enough. You have to look for anomalies at every single part of the layer across a lot of metrics.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1164\" target=\"_blank\" rel=\"noreferrer noopener\">19.24<\/a>:<strong> So your position is we can use the existing tools. We may have to log more things.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1173\" target=\"_blank\" rel=\"noreferrer noopener\">19.33<\/a>: We may have to log more things, and then start building simple ML models to be able to do anomaly detection.\u00a0<\/p>\n<p>Think of managing any machine, any LLM model, any agent as really like a fraud detection pipeline. So every single time you\u2019re looking for \u201cWhat are the simplest signs of fraud?\u201d And that can happen across various factors. But we need more logging. And again you don\u2019t need external tools for that. You can set up your own loggers as well.<\/p>\n<p>Most of the people I know have been setting up their own loggers within their companies. So you can simply use telemetry to be able to a.) define a set and use the general logs, and b.) be able to define your own custom logs as well, depending on your agent pipeline itself. You can define \u201cThis is what it\u2019s trying to do\u201d and log more things across those things, and then start building small machine learning models to look for what\u2019s going on over there.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1236\" target=\"_blank\" rel=\"noreferrer noopener\">20.36<\/a>: <strong>So what is the state of \u201cWhere we are? How many teams are doing this?\u201d<\/strong>\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1242\" target=\"_blank\" rel=\"noreferrer noopener\">20.42<\/a>: Very few. Very, very few. Maybe just the top bits. The ones who are doing reinforcement learning training and using RL environments, because that\u2019s where they\u2019re getting their data to do RL. But people who are not using RL to be able to retrain their model, they\u2019re not really doing much of this part; they\u2019re still depending very much on external accounts.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1272\" target=\"_blank\" rel=\"noreferrer noopener\">21.12<\/a>: <strong>I\u2019ll get back to RL in a second. But one topic you raised when you pointed out the transition from MLOps to LLMOps was the importance of FinOps, which is, for our listeners, basically managing your cloud computing costs\u2014or in this case, increasingly mastering token economics. Because basically, it\u2019s one of these things that I think can bite you.<\/strong><\/p>\n<p><strong>For example, the first time you use Claude Code, you go, \u201cOh, man, this tool is powerful.\u201d And then boom, you get an email with a bill. I see, that\u2019s why it\u2019s powerful. And you multiply that across the board to teams who are starting to maybe deploy some of these things. And you see the importance of FinOps.<\/strong><\/p>\n<p><strong>So where are we, Abi, as far as tooling for FinOps in the age of generative AI and also the practice of FinOps in the age of generative AI?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1339\" target=\"_blank\" rel=\"noreferrer noopener\">22.19<\/a>: Less than 5%, maybe even 2% of the way there.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1344\" target=\"_blank\" rel=\"noreferrer noopener\">22:24<\/a>: <strong>Really? But obviously everyone\u2019s aware of it, right? Because at some point, when you deploy, you become aware.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1353\" target=\"_blank\" rel=\"noreferrer noopener\">22.33<\/a>: Not enough people. A lot of people just think about FinOps as cloud, basically the cloud cost. And there are different kinds of costs in the cloud. One of the things people are not doing enough is not profiling their models properly, which is [determining] \u201cWhere are the costs really coming from? Our models\u2019 compute power? Are they taking too much RAM?\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1378\" target=\"_blank\" rel=\"noreferrer noopener\">22.58<\/a>: <strong>Or are we using reasoning when we don\u2019t need it?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1380\" target=\"_blank\" rel=\"noreferrer noopener\">23.00<\/a>: Exactly. Now that\u2019s a problem we solve very differently. That\u2019s where yes, you can do kernel fusion. Define your own custom kernels. Right now there\u2019s a massive number of people who think we need to rewrite kernels for everything. It\u2019s only going to solve one problem, which is the compute-bound problem. But it\u2019s not going to solve the memory-bound problem. Your data engineering pipelines aren\u2019t what\u2019s going to solve your memory-bound problems.<\/p>\n<p>And that\u2019s where most of the focus is missing. I\u2019ve mentioned it in the book as well: Data engineering is the foundation of first being able to solve the problems. And then we moved to the compute-bound problems. Do not start optimizing the kernels over there. And then the third part would be the communication-bound problem, which is \u201cHow do we make these GPUs talk smarter with each other? How do we figure out the agent consensus and all of those problems?\u201d<\/p>\n<p>Now that\u2019s a communication problem. And that\u2019s what happens when there are different levels of bandwidth. Everybody\u2019s dealing with the internet bandwidth as well, the kind of serving speed as well, different kinds of cost and every kind of transitioning from one node to another. If we\u2019re not really hosting our own infrastructure, then that\u2019s a different problem, because it depends on \u201cWhich server do you get assigned your GPUs on again?\u201d<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1460\" target=\"_blank\" rel=\"noreferrer noopener\">24.20<\/a>: <strong>Yeah, yeah, yeah. I want to give a shout out to Ray\u2014I\u2019m an advisor to Anyscale\u2014because Ray basically is built for these sorts of pipelines because it can do fine-grained utilization and help you decide between CPU and GPU. And just generally, you don\u2019t think that the teams are taking token economics seriously?<\/strong><\/p>\n<p><strong>I guess not. How many people have I heard talking about caching, for example? Because if it\u2019s a prompt that [has been] answered before, why do you have to go through it again?\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1507\" target=\"_blank\" rel=\"noreferrer noopener\">25.07<\/a>: I think plenty of people have started implementing KV caching, but they don\u2019t really know.\u00a0.\u00a0. Again, one of the questions people don\u2019t understand is \u201cHow much do we need to store in the memory itself, and how much do we need to store in the cache?\u201d which is the big memory question. So that\u2019s the one I don\u2019t think people are able to solve. A lot of people are storing too much stuff in the cache that should actually be stored in the RAM itself, in the memory.<\/p>\n<p>And there are generalist applications that don\u2019t really understand that this agent doesn\u2019t really need access to the memory. There\u2019s no point. It\u2019s just lost in the throughput really. So I think the problem isn\u2019t really caching. The problem is that differentiation of understanding for people.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1555\" target=\"_blank\" rel=\"noreferrer noopener\">25.55<\/a>:<strong> Yeah, yeah, I just threw that out as one element. Because obviously there\u2019s many, many things to mastering token economics. So you, you brought up reinforcement learning. A few years ago, obviously people got really into \u201cLet\u2019s do fine-tuning.\u201d But then they quickly realized.\u00a0.\u00a0. And actually fine-tuning became easy because basically there became so many services where you can just focus on labeled data. You upload your labeled data, boom, come back from lunch, you have a fine-tuned model.<\/strong><\/p>\n<p><strong>But then people realize that \u201cI fine-tuned, but the model that results isn\u2019t really as good as my fine-tuning data.\u201d And then obviously RAG and context engineering came into the picture. Now it seems like more people are again talking about reinforcement learning, but in the context of LLMs. And there\u2019s a lot of libraries, many of them built on Ray, for example. But it seems like what\u2019s missing, Abi, is that fine-tuning got to the point where I can sit down a domain expert and say, \u201cProduce labeled data.\u201d And basically the domain expert is a first-class participant in fine-tuning.<\/strong><\/p>\n<p><strong>As best I can tell, for reinforcement learning, the tools aren\u2019t there yet. The UX hasn\u2019t been figured out in order to bring in the domain experts as the first-class citizen in the reinforcement learning process\u2014which they need to be because a lot of the stuff really resides in their brain.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1665\" target=\"_blank\" rel=\"noreferrer noopener\">27.45<\/a>: The big problem here, and very, very much to the point of what you pointed out, is the tools aren\u2019t really there. And one very specific thing I can tell you is most of the reinforcement learning environments that you\u2019re seeing are static environments. Agents are not learning statically. They are learning dynamically. If your RL environment cannot adapt dynamically, which basically in 2018, 2019, emerged as the OpenAI Gym and a lot of reinforcement learning libraries were coming out.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1698\" target=\"_blank\" rel=\"noreferrer noopener\">28.18<\/a>: There is a line of work called curriculum learning, which is basically adapting your model\u2019s difficulty to the results itself. So basically now that can be used in reinforcement learning, but I\u2019ve not seen any practical implementation of using curriculum learning for reinforcement learning environments. So people create these environments\u2014fantastic. They work well for a little bit of time, and then they become useless.<\/p>\n<p>So that\u2019s where even OpenAI, Anthropic, those companies are struggling as well. They\u2019ve paid heavily in contracts, which are yearlong contracts to say, \u201cCan you build this vertical environment? Can you build that vertical environment?\u201d and that works fantastically But once the model learns on it, then there\u2019s nothing else to learn. And then you go back into the question of, \u201cIs this data fresh? Is this adaptive with the world?\u201d And it becomes the same RAG problem over again.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1758\" target=\"_blank\" rel=\"noreferrer noopener\">29.18<\/a>: <strong>So maybe the problem is with RL itself. Maybe maybe we need a different paradigm. It\u2019s just too hard.\u00a0<\/strong><\/p>\n<p><strong>Let me close by looking to the future. The first thing is\u2014the space is moving so hard, this might be an impossible question to ask, but if you look at, let\u2019s say, 6 to 18 months, what are some things in the research domain that you think are not being talked enough about that might produce enough practical utility that we will start hearing about them in 6 to 12, 6 to 18 months?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1795\" target=\"_blank\" rel=\"noreferrer noopener\">29.55<\/a>: One is how to profile your machine learning models, like the entire systems end-to-end. A lot of people do not understand them as systems, but only as models. So that\u2019s one thing which will make a massive amount of difference. There are a lot of AI engineers today, but we don\u2019t have enough system design engineers.<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1816\" target=\"_blank\" rel=\"noreferrer noopener\">30.16<\/a>: <strong>This is something that Ion Stoica at Sky Computing Lab has been giving keynotes about. Yeah. Interesting.\u00a0<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1823\" target=\"_blank\" rel=\"noreferrer noopener\">30.23<\/a>: The second part is.\u00a0.\u00a0. I\u2019m optimistic about seeing curriculum learning applied to reinforcement learning as well, where our RL environments can adapt in real time so when we train agents on them, they are dynamically adapting as well. That\u2019s also [some] of the work being done by labs like Circana, which are working in artificial labs, artificial light frame, all of that stuff\u2014evolution of any kind of machine learning model accuracy.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1857\" target=\"_blank\" rel=\"noreferrer noopener\">30.57<\/a>: The third thing where I feel like the communities are falling behind massively is on the data engineering side. That\u2019s where we have massive gains to get.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1869\" target=\"_blank\" rel=\"noreferrer noopener\">31.09<\/a>: <strong>So on the data engineering side, I\u2019m happy to say that I advise several companies in the space that are completely focused on tools for these new workloads and these new data types.\u00a0<\/strong><\/p>\n<p><strong>Last question for our listeners: What mindset shift or what skill do they need to pick up in order to position themselves in their career for the next 18 to 24 months?<\/strong><\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1900\" target=\"_blank\" rel=\"noreferrer noopener\">31.40<\/a>: For anybody who\u2019s an AI engineer, a machine learning engineer, an LLMOps engineer, or an MLOps engineer, first learn how to profile your models. Start picking up Ray very quickly as a tool to just get started on, to see how distributed systems work. You can pick the LLM if you want, but start understanding distributed systems first. And once you start understanding those systems, then start looking back into the models itself.\u00a0<\/p>\n<p><a href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_Abi_Aryan.mp3#t=1931\" target=\"_blank\" rel=\"noreferrer noopener\">32.11<\/a>: <strong>And with that, thank you, Abi.<\/strong><\/p>\n<\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>Generative AI in the Real World Generative AI in the Real World: The LLMOps Shift with Abi Aryan Play Episode Pause Episode Mute\/Unmute Episode Rewind 10 Seconds 1x Fast Forward 30 seconds 00:00 \/ 32m 16s Subscribe Share MLOps is dead. Well, not really, but for many the job is evolving into LLMOps. In this [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":311555,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[141427,12386],"tags":[81777,141429,141428,108637,11637],"dealstore":[],"offerexpiration":[],"class_list":["post-311554","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-generative-ai-in-the-real-world","category-podcast","tag-abi","tag-aryan","tag-llmops","tag-oreilly","tag-shift"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=311554\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Generative AI in the Real World Generative AI in the Real World: The LLMOps Shift with Abi Aryan Play Episode Pause Episode Mute\/Unmute Episode Rewind 10 Seconds 1x Fast Forward 30 seconds 00:00 \/ 32m 16s Subscribe Share MLOps is dead. Well, not really, but for many the job is evolving into LLMOps. In this [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=311554\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-11-22T16:03:19+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1600\" \/>\n\t<meta property=\"og:image:height\" content=\"1600\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"25 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=311554#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly\",\"datePublished\":\"2025-11-22T16:03:19+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554\"},\"wordCount\":4989,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png\",\"keywords\":[\"Abi\",\"Aryan\",\"LLMOps\",\"OReilly\",\"shift\"],\"articleSection\":[\"Generative AI in the Real World\",\"Podcast\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=311554#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=311554\",\"url\":\"https:\/\/fivemor.com\/?p=311554\",\"name\":\"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png\",\"datePublished\":\"2025-11-22T16:03:19+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=311554#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=311554\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=311554#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png\",\"width\":1600,\"height\":1600},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=311554#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=311554","og_locale":"en_US","og_type":"article","og_title":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network","og_description":"Generative AI in the Real World Generative AI in the Real World: The LLMOps Shift with Abi Aryan Play Episode Pause Episode Mute\/Unmute Episode Rewind 10 Seconds 1x Fast Forward 30 seconds 00:00 \/ 32m 16s Subscribe Share MLOps is dead. Well, not really, but for many the job is evolving into LLMOps. In this [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=311554","og_site_name":"Som2ny Network","article_published_time":"2025-11-22T16:03:19+00:00","og_image":[{"width":1600,"height":1600,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"25 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=311554#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=311554"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly","datePublished":"2025-11-22T16:03:19+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=311554"},"wordCount":4989,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=311554#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png","keywords":["Abi","Aryan","LLMOps","OReilly","shift"],"articleSection":["Generative AI in the Real World","Podcast"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=311554#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=311554","url":"https:\/\/fivemor.com\/?p=311554","name":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=311554#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=311554#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png","datePublished":"2025-11-22T16:03:19+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=311554#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=311554"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=311554#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/11\/Podcast_Cover_GenAI_in_the_Real_World-1600x1600.png","width":1600,"height":1600},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=311554#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"The LLMOps Shift with Abi Aryan \u2013 O\u2019Reilly"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/311554","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=311554"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/311554\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/311555"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=311554"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=311554"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=311554"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=311554"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=311554"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}