{"id":283504,"date":"2025-06-09T11:28:52","date_gmt":"2025-06-09T11:28:52","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/ml-model-serving-with-fastapi-and-redis-for-faster-predictions\/"},"modified":"2025-06-09T11:28:52","modified_gmt":"2025-06-09T11:28:52","slug":"ml-model-serving-with-fastapi-and-redis-for-faster-predictions","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=283504","title":{"rendered":"ML Model Serving with FastAPI and Redis for faster predictions"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p>Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Users, on the other hand, expect instant feedback. That\u2019s where latency becomes a real problem. Technically speaking, one of the biggest problems is redundant computation when the same input triggers the same slow process repeatedly. In this blog, I\u2019ll show you how to fix that. We will build a FastAPI-based ML service and integrate Redis caching to return repeated predictions in milliseconds.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-fastapi\">What is FastAPI?<\/h2>\n<p><a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2020\/11\/fastapi-the-right-replacement-for-flask\/\" target=\"_blank\" rel=\"noreferrer noopener\">FastAPI<\/a> is a modern, high-performance web framework for building APIs with Python. It uses <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2021\/05\/introduction-to-python-programming-beginners-guide\/\" target=\"_blank\" rel=\"noreferrer noopener\">Python<\/a>\u2018s type hints for data validation and automatic generation of interactive API documentation using Swagger UI and ReDoc. Built on top of Starlette and Pydantic, FastAPI supports asynchronous programming, making it comparable in performance to Node.js and Go. Its design facilitates rapid development of robust, production-ready APIs, making it an excellent choice for deploying machine learning models as scalable RESTful services.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-what-is-redis\">What is Redis?<\/h2>\n<p>Redis (Remote Dictionary Server) is an open-source, in-memory data structure store that functions as a database, cache, and message broker. By storing data in memory, Redis offers ultra-low latency for read and write operations, making it ideal for caching frequent or computationally intensive tasks like machine learning model predictions. It supports various data structures, including strings, lists, sets, and hashes, and provides features like key expiration (TTL) for efficient cache management.<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-why-combine-fastapi-and-redis\">Why Combine FastAPI and Redis?<\/h2>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter\"><img decoding=\"async\" src=\"https:\/\/lh7-rt.googleusercontent.com\/docsz\/AD_4nXfroMuCYanydro6m6a1R0hDefVmv2MhUEEw9QJuz-kkogTVYIjcL3BoYIrCm_s6geOXKF7R_z9fDFlM0f-DvzpZxf-PozNlQxchcxNJ8o_MaffnXhQY1UF_PdhRTVNk4zUUiQFXSg?key=_kgsNweIOVxq2e6ywKmZ1g\" alt=\"\"\/><\/figure>\n<\/div>\n<p>Integrating <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2020\/11\/fastapi-the-right-replacement-for-flask\/\" target=\"_blank\" rel=\"noreferrer noopener\">FastAPI<\/a> with Redis creates a system that is both responsive and efficient. FastAPI serves as a swift and reliable interface for handling API requests, while Redis acts as a caching layer to store the results of previous computations. When the same input is received again, the result can be retrieved instantly from Redis, bypassing the need for recomputation. This approach reduces latency, lowers computational load, and enhances the scalability of your application. In distributed environments, Redis serves as a centralised cache accessible by multiple FastAPI instances, making it an excellent fit for production-grade machine learning deployments.<\/p>\n<p>Now, let\u2019s walk through the implementation of a FastAPI application that serves machine learning model predictions with Redis caching. This setup ensures that repeated requests with the same input are served quickly from the cache, reducing computation time and improving response times. The steps are mentioned below:\u00a0<\/p>\n<ol class=\"wp-block-list\">\n<li>Loading a Pre-trained Model<\/li>\n<li>Creating a FastAPI Endpoint for Predictions<\/li>\n<li>Setting Up Redis Caching<\/li>\n<li>Measuring Performance Gains<\/li>\n<\/ol>\n<p>Now, let\u2019s see these steps in more detail.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-1-loading-a-pre-trained-model\">Step 1: Loading a Pre-trained Model<\/h3>\n<p>First, assume that you already have a trained machine learning model that is ready to deploy. In practice, most of the models are trained offline (like a scikit-learn model, a <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2022\/03\/a-basic-introduction-to-tensorflow-in-deep-learning\/\" target=\"_blank\" rel=\"noreferrer noopener\">TensorFlow<\/a>\/Pytorch model, etc), saved to disk, and then loaded into a serving app. For our example, we will create a simple scikit-learn classifier that will be trained on the famous Iris flower dataset and saved using joblib. If you already have a saved model file, you can skip the training part and just load it. Here\u2019s how to train a model and then load it for serving:<\/p>\n<pre class=\"wp-block-code\"><code>from sklearn.datasets import load_iris\nfrom sklearn.ensemble import RandomForestClassifier\nimport joblib\n\n# Load example dataset and train a simple model (Iris classification)\nX, y = load_iris(return_X_y=True)\n\n# Train the model\nmodel = RandomForestClassifier().fit(X, y)\n\n# Save the trained model to disk\njoblib.dump(model, \"model.joblib\")\n\n# Load the pre-trained model from disk (using the saved file)\nmodel = joblib.load(\"model.joblib\")\n\nprint(\"Model loaded and ready to serve predictions.\")<\/code><\/pre>\n<p>In the above code, we have used scikit-learn\u2019s built-in <a href=\"https:\/\/www.kaggle.com\/datasets\/uciml\/iris\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">Iris dataset<\/a>, trained a random forest classifier on it, and then saved that model to a file called <em>model.joblib<\/em>. After that, we have loaded it back using joblib.load. The joblib library is pretty common when it comes to saving scikit-learn models, mostly because it is good at handling NumPy arrays inside models. After this step, we have a model object ready to predict on new data. Just a heads-up, though, you can use any pre-trained model here, the way you serve it using FastAPI, and also cached results would be more or less the same. The only thing is, the model should have a predict method that takes in some input and produces the result. Also, make sure that the model\u2019s prediction stays the same every time you give it the same input (so it\u2019s deterministic). If it\u2019s not, caching would be problematic for non-deterministic models as it would return incorrect results.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-2-creating-a-fastapi-prediction-endpoint\">Step 2: Creating a FastAPI Prediction Endpoint<\/h3>\n<p>Now that we have a model, let\u2019s use it via API. We will be using FASTAPI to create a web server that attends to prediction requests. FASTAPI makes it easy to define an endpoint and map request parameters to Python function arguments. In our example, we will assume the model accepts four features. And will create a GET endpoint <code>\/predict<\/code> that accepts these features as query parameters and returns the model\u2019s prediction.<\/p>\n<pre class=\"wp-block-code\"><code>from fastapi import FastAPI\nimport joblib\n\napp = FastAPI()\n\n# Load the trained model at startup (to avoid re-loading on every request)\nmodel = joblib.load(\"model.joblib\")  # Ensure this file exists from the training step\n\n@app.get(\"\/predict\")\ndef predict(sepal_length: float, sepal_width: float, petal_length: float, petal_width: float):\n    \"\"\" Predict the Iris flower species from input measurements. \"\"\"\n    \n    # Prepare the features for the model as a 2D list (model expects shape [n_samples, n_features])\n    features = [[sepal_length, sepal_width, petal_length, petal_width]]\n    \n    # Get the prediction (in the iris dataset, prediction is an integer class label 0,1,2 representing the species)\n    prediction = model.predict(features)[0]  # Get the first (only) prediction\n    \n    return {\"prediction\": str(prediction)}\n<\/code><\/pre>\n<p>In the above code, we have made a FastAPI app, and upon executing the file, it starts the API server. FastAPI is super fast for Python, so it can handle lots of requests easily. Then we load the model just at the start because doing it again and again on every request would be slow, so we keep it in memory, which is ready to use. We created a <code>\/predict<\/code> endpoint with <code>@app.get<\/code>, GET makes testing easy since we can just pass things in the URL, but in real projects, you will probably want to use POST, especially if sending big or complex input like images or JSON. This function takes 4 inputs: <code>sepal_length<\/code>, <code>sepal_width<\/code>, <code>petal_length<\/code>, and <code>petal_width<\/code>, and FastAPI auto reads them from the URL. Inside the function, we put all the inputs into a 2D list (because scikit-learn accepts only a 2D array), then we call <code>model.predict()<\/code>, and it gives us a list. Then we return it as JSON like<code> { \u201cprediction\u201d: \u201c...\u201d}<\/code>.<\/p>\n<p>Therefore, now it works, you can run it using <code>uvicorn main:app --reload<\/code>, hit <code>\/predict<\/code>, endpoint and get results. Even if you send the same input again, it still runs the model again, which is not good, so the next step is adding Redis to cache the previous results and skip redoing them.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-3-adding-redis-caching-for-predictions\">Step 3: Adding Redis Caching for Predictions<\/h3>\n<p>To cache the model output, we will be using Redis. First, make sure the Redis server is running. You can install it locally or just run a Docker container; it usually runs on port <strong>6379<\/strong> by default. We will be using the Python redis library to talk to the server.<\/p>\n<p>So the idea is simple: when a request comes in, create a unique key that represents the input. Then check if the key exists in Redis; if that key is already there, which means we already cached this before, so we just return the saved result, no need to call the model again. If not there, we do <code>model.predict<\/code>, get the output, save it in Redis, and send back the prediction.<\/p>\n<p>Let\u2019s now update the FastAPI app to add this cache logic.<\/p>\n<pre class=\"wp-block-code\"><code>!pip install redis\nimport redis  # New import to use Redis\n\n# Connect to a local Redis server (adjust host\/port if needed)\ncache = redis.Redis(host=\"localhost\", port=6379, db=0)\n\n@app.get(\"\/predict\")\ndef predict(sepal_length: float, sepal_width: float, petal_length: float, petal_width: float):\n    \"\"\"\n    Predict the species, with caching to speed up repeated predictions.\n    \"\"\"\n    # 1. Create a unique cache key from input parameters\n    cache_key = f\"{sepal_length}:{sepal_width}:{petal_length}:{petal_width}\"\n    \n    # 2. Check if the result is already cached in Redis\n    cached_val = cache.get(cache_key)\n    \n    if cached_val:\n        # If cache hit, decode the bytes to a string and return the cached prediction\n        return {\"prediction\": cached_val.decode(\"utf-8\")}\n    \n    # 3. If not cached, compute the prediction using the model\n    features = [[sepal_length, sepal_width, petal_length, petal_width]]\n    prediction = model.predict(features)[0]\n    \n    # 4. Store the result in Redis for next time (as a string)\n    cache.set(cache_key, str(prediction))\n    \n    # 5. Return the freshly computed prediction\n    return {\"prediction\": str(prediction)}<\/code><\/pre>\n<p>In the above code, we added Redis now. First, we made a client using <code>redis.Redis()<\/code>. It connects to the Redis server. Using db=0 by default. Then we created a cache key just by joining the input values. Here it works because the inputs are simple numbers, but for complex ones it\u2019s better to use a hash or a JSON string. The key must be unique for each input. We have used <code>cache.get(cache_key)<\/code>. If it finds the same key, it returns that, which makes it fast, and with this, there is no need to rerun the model. But if it is not found in the cache, we need to run the model and get the prediction. Finally, save that in Redis using <code>cache.set()<\/code>. So next time, when the same input comes, it\u2019s already there, and caching would be fast.<\/p>\n<h3 class=\"wp-block-heading\" id=\"h-step-4-testing-and-measuring-performance-gains\">Step 4: Testing and Measuring Performance Gains<\/h3>\n<p>Now that our FastAPI app is running and is connected to Redis, it\u2019s time for us to test how caching improves the response time. Here, I will demonstrate how to use Python\u2019s requests library to call the API twice with the same input and measure the time taken for each call. Also, make sure that you start your FastAPI before running the test code:<\/p>\n<pre class=\"wp-block-code\"><code>import requests, time\n# Sample input to predict (same input will be used twice to test caching)\nparams = {\n\"sepal_length\": 5.1,\n\"sepal_width\": 3.5,\n\"petal_length\": 1.4,\n\"petal_width\": 0.2\n}\n\n# First request (expected to be a cache miss, will run the model)\nstart = time.time()\nresponse1 = requests.get(\"http:\/\/localhost:8000\/predict\", params=params)\nelapsed1 = time.time() - start\nprint(\"First response:\", response1.json(), f\"(Time: {elapsed1:.4f} seconds)\")<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"791\" height=\"54\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP1.webp\" alt=\"Output 1\" class=\"wp-image-237030\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP1.webp 791w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP1-300x20.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP1-768x52.webp 768w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP1-150x10.webp 150w\" sizes=\"auto, (max-width: 791px) 100vw, 791px\"\/><\/figure>\n<\/div>\n<pre class=\"wp-block-code\"><code># Second request (same params, expected cache hit, no model computation)\nstart = time.time()\nresponse2 = requests.get(\"http:\/\/localhost:8000\/predict\", params=params)\nelapsed2 = time.time() - start\nprint(\"Second response:\", response2.json(), f\"(Time: {elapsed2:.6f}seconds)\")<\/code><\/pre>\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"735\" height=\"49\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP2.webp\" alt=\"Output 2\" class=\"wp-image-237031\" style=\"object-fit:cover\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP2.webp 735w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP2-300x20.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/OP2-150x10.webp 150w\" sizes=\"auto, (max-width: 735px) 100vw, 735px\"\/><\/figure>\n<\/div>\n<p>When you run this, you should see the first request return a result. Then the second request returns the same result, but noticeably faster. For example, you might find the first call took on the order of tens of milliseconds (depending on model complexity), while the second call might be a few milliseconds or less. In our simple demo with a lightweight model, the difference might be small (since the model itself is fast), but the effect is drastic for heavier models. <\/p>\n<h3 class=\"wp-block-heading\" id=\"h-comparison\">Comparison<\/h3>\n<p>To put this into perspective, let\u2019s consider what we achieved:<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Without caching:<\/strong> Every request, even identical ones, would hit the model. If the model takes 100 ms per prediction, 10 identical requests would collectively still take ~1000 ms.<\/li>\n<li><strong>With caching:<\/strong> The first request takes the full hit (100 ms), but the next 9 identical requests might take, say, 1\u20132 ms each (just a Redis lookup and returning data). So those 10 requests might total ~120 ms instead of 1000 ms, a ~8x speed-up in this scenario.\u00a0<\/li>\n<\/ul>\n<p>In real experiments, caching can lead to order-of-magnitude improvements. In e-commerce, for example, <em>using Redis meant returning recommendations in microseconds for repeat requests, versus<\/em> <em>having to recompute them with the full model serve pipeline<\/em>. The performance gain will depend on how expensive your model inference is. The more complex the model, the more you benefit from caching on repeated calls. It also depends on request patterns: if every request is unique, the cache won\u2019t help (no repeats to serve from memory), but many applications do see overlapping requests (e.g., popular search queries, recommended items, etc.).<\/p>\n<p>You can also check your Redis cache directly to verify it\u2019s storing keys.\u00a0<\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\">Conclusion<\/h2>\n<p>In this blog, we demonstrated how FastAPI and Redis can work in collaboration to accelerate ML model serving. FastAPI provides a fast and easy-to-build API layer for serving predictions, and Redis adds a caching layer that significantly reduces latency and CPU load for repeated computations. By avoiding repeated model calls, we have improved responsiveness and also enabled the system to handle more requests with the same resources.\u00a0<\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/janvikumari01\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_ToTu2tx.webp\" width=\"48\" height=\"48\" alt=\"Janvi Kumari\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Hi, I am Janvi, a passionate data science enthusiast currently working at Analytics Vidhya. My journey into the world of data began with a deep curiosity about how we can extract meaningful insights from complex datasets.<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Users, on the other hand, expect instant feedback. That\u2019s where latency becomes a real problem. Technically speaking, one of the biggest problems is [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":283505,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[113073,11027,1168,6638,113074,7416],"dealstore":[],"offerexpiration":[],"class_list":["post-283504","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-fastapi","tag-faster","tag-model","tag-predictions","tag-redis","tag-serving"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=283504\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Users, on the other hand, expect instant feedback. That\u2019s where latency becomes a real problem. Technically speaking, one of the biggest problems is [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=283504\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-06-09T11:28:52+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=283504#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"ML Model Serving with FastAPI and Redis for faster predictions\",\"datePublished\":\"2025-06-09T11:28:52+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504\"},\"wordCount\":1724,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp\",\"keywords\":[\"FastAPI\",\"Faster\",\"Model\",\"Predictions\",\"Redis\",\"Serving\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=283504#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=283504\",\"url\":\"https:\/\/fivemor.com\/?p=283504\",\"name\":\"ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp\",\"datePublished\":\"2025-06-09T11:28:52+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=283504#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=283504\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=283504#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=283504#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"ML Model Serving with FastAPI and Redis for faster predictions\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=283504","og_locale":"en_US","og_type":"article","og_title":"ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network","og_description":"Ever waited too long for a model to return predictions? We have all been there. Machine learning models, especially the large, complex ones, can be painfully slow to serve in real time. Users, on the other hand, expect instant feedback. That\u2019s where latency becomes a real problem. Technically speaking, one of the biggest problems is [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=283504","og_site_name":"Som2ny Network","article_published_time":"2025-06-09T11:28:52+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=283504#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=283504"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"ML Model Serving with FastAPI and Redis for faster predictions","datePublished":"2025-06-09T11:28:52+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=283504"},"wordCount":1724,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=283504#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp","keywords":["FastAPI","Faster","Model","Predictions","Redis","Serving"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=283504#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=283504","url":"https:\/\/fivemor.com\/?p=283504","name":"ML Model Serving with FastAPI and Redis for faster predictions - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=283504#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=283504#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp","datePublished":"2025-06-09T11:28:52+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=283504#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=283504"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=283504#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Accelerate-Machine-Learning-Model-Serving-with-FastAPI-and-Redis-Caching_.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=283504#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"ML Model Serving with FastAPI and Redis for faster predictions"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/283504","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=283504"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/283504\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/283505"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=283504"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=283504"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=283504"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=283504"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=283504"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}