Tech ONTAP Podcast Episode 406 – AI Powered NetApp Docs


Episode 406 of the Tech ONTAP Podcast takes us behind the scenes of NetApp Docs, the company’s most visited website, where documentation has evolved from static files into an intelligent, interactive, and multilingual platform.

Key Takeaways:

  1. From Walled Gardens to Openness
    • NetApp moved from proprietary XML systems to GitHub and open standards, making all product documentation freely available online and vastly improving searchability.
  2. AI-Powered Assistance
    • The team built Doc the Owl, a generative AI chatbot trained on more than 120,000 documents.
    • Using retrieval-augmented generation (RAG), Doc delivers accurate answers with citations while continuously learning from customer feedback.
  3. Continuous Innovation
    • Daily content builds in 9 languages, accelerated by Kubernetes, Ansible, Trident, and GitHub Actions.
    • Integration with cloud providers (AWS, Azure, Google) ensures customers can access authoritative information across ecosystems.
  4. Future Directions
    • Embedding AI-powered search directly into products like ONTAP System Manager and BlueXP.
    • Leveraging AutoSupport data for best-practice recommendations.
    • Expanding machine translation quality with AI oversight.

Why It Matters
NetApp Docs is no longer just documentation—it’s a data-driven ecosystem that improves customer experience, empowers support teams, and sets the stage for AI-driven knowledge delivery.

Finding the podcast

Check it out here, like and subscribe and all that jazz (now hosted on the NetApp YouTube channel!):

If you prefer audio only, I also still offer that:

You can also find the Tech ONTAP Podcast on:

Transcription

The following transcript was generated using Descript’s speech to text service and then further edited. As it is AI generated, YMMV.

Episode 406 – NetApp Docs
===

[00:00:00]

Justin Parisi: All right. I am here in my studio, in my house, and I have a whole bunch of people here today to talk to us about NetApp docs which is one of our things we offer here at NetApp to help enhance your overall experience. And we’ll get more into that later. But first I want to go through the room here and introduce everybody.

So we’ll start with Aksel Davis, since he’s the first person listed here. Aksel, what do you do here at NetApp and how do we reach you?

Aksel Davis: Hey, thanks, Justin. I’m a software engineering manager here in information engineering at NetApp. So I run all of our supporting services for our documentation program.

And we also run our content automation program, which works with our engineering teams to produce our API and CLI and other autogenerated documentation.

Justin Parisi: Okay, excellent. Also with us, Grant Glass. Grant, what do you do here at NetApp and how do I reach you?

Grant Glass: So I work as a applied and data scientist at NetApp.

So I work a lot in generative AI and governance and I’m [00:01:00] reachable at NetApp at the easiest possible email address [email protected]. I’m very happy to be here and excited to talk about our product documentation.

Justin Parisi: So you’re a data scientist, but I don’t see any bunson burners or glass beakers behind you.

What’s going on there? What’s up with that?

Grant Glass: Well, data doesn’t, you know, we have our spreadsheets and databases and arrays and python notebooks, right. That are all on the computer in front of me. So, the magic really happens behind the screen as opposed to behind it. All right.

Justin Parisi: So, you mentioned Gen AI and we have the original Jen AI here, Jen Kaufman.

Jen, what do you do here at NetApp? How do I reach you?

Jen Kaufman: Well very much the way Grant is a data scientist with no science. I’m an information engineer without engineering. I’m the director of the Information engineering [00:02:00] organization at NetApp. That’s our fancy terminology for tech pubs. So all the content that is generated to help customers and employees and partners use NetApp products successfully. We have a small but mighty team of content creators and Aksel Davis’s development team working to make sure that we can have everyone use our products successfully out there in the world.

Justin Parisi: Alright, and last but not least, Adam Newton. Adam, what do you do here at NetApp? How do we reach you?

Adam Newton: Currently the most important thing I’m doing is trying to prevent my cat from coming across to the screen here.

So I’m holding my hand out. My name’s Adam Newton. I’m the senior director of our Global Content Experience Services team, which comprises Jennifer’s team, the information engineering team, pubs and globalization. We handle all of the product and collateral globalization needs for the company and also all the data scientists who are helping us on the new frontiers of generative ai.

Justin Parisi: Alright, so [00:03:00] we’re here to talk about NetApp docs and I don’t know if you’re familiar with the documentation of NetApp or the history of the documentation, but it has come a long way from previous times.

I mean, I’ve been here a long time. I’ve seen the documentation in previous iterations and I could say confidently that’s gotten better because you could actually find stuff now, right? I mean, before it was tough. We’d shoved things in PDFs, which aren’t searchable. Web searches wouldn’t find things, so we’re making a lot of progress in that regard.

But NetApp docs is a little bit different, I think, than the overall documentation effort at NetApp. So Adam, tell us what NetApp docs itself is.

Adam Newton: So I’d actually like to start by going backwards a little bit and just say, there were some guiding principles that informed where we are today. Not the least of which was that we realized several years ago, primarily with the advent of the cloud portfolio at NetApp, that our bespoke system that we were using that was sort of a walled garden which we [00:04:00] authored content was a private domain that we weren’t really able to move the business, nor were we able to invite people in to help us produce documentation.

So we said, Hey, we’re gonna go to a more open. Mindset and accordingly our systems and processes followed. And Aksel here on this call was pivotal in helping us move from that closed XML based system to GitHub and open source standards. And so that was the mindset that led us to where we are today.

And then another big thing that we did you alluded to it Justin, about findability, searchability was all of our content is freely available on the internet. And that was a huge change for us and has once again, I think been part of our openness. And it has led to a lot of the innovation we’ve had in the last two, three years. So we are a GitHub based organization called NetApp Docs. The site itself, docs.netapp.com is by far [00:05:00] and away the most used website at NetApp. The most visited website. And we have content through build automation, being continuously built in nine languages.

So it’s quite a powerhouse and we see it’s creating value in the pre-sales context. Let’s say Justin Parisi was a skeptical prospect customer.

Justin Parisi: I am, I am a skeptical prospect.

Adam Newton: I can’t believe Justin Parisi ever would be, but let’s say for the sake of argument, he were that person. We do see people using the site to really learn about the products before they’ve bought them, but also to use the products after they’ve bought them. And we see a huge constituency also within NetApp. Support and other organizations extensively use our documentation.

I’ll stop there. Could go on for a while, but hopefully that’s helpful.

Justin Parisi: That is helpful. Yeah, absolutely. So Jen, you’ve been involved with information engineering for a long time. I think you started out as

Jen Kaufman: 20 years.

Justin Parisi: Sorry. Someone who would write things, right? So now, you’ve moved up the ranks here [00:06:00] and you are up at the upper echelons of these things. Now, what did you envision NetApp docs being, what sort of use cases were you seeing it materialize as and then how has it changed? Is it any different? Did it fit what you thought it was gonna be or is there more to it?

Jen Kaufman: Oh, that’s a really good question. It’s almost like you interview people. I’ve done this before all the time, and you kind of know what you’re, I mean, I’ll say this, that the purpose of technical documentation or of technical content, it has not changed.

The purpose is to make it easy for people to use products that are complicated. So you really do need information that will help people be successful. And this is employees, partners, customers, lots of different kinds of people who come to our site, docs.netapp.com, or other places looking for information.

Our content is used by people to create more content, right? So folks like you even might take some [00:07:00] content from our site, repurpose it and help somebody else, or a group of other people to be successful with the products. So that mission, to help make sure our products can be used and sold, has not changed. The way we do it changes all the time, right?

The technology changes, the expectations and needs of the customers change the landscape around us changes how quickly we need to do things, changes, the surfaces that customers expect to consume this information from change, which kind of brings us to our topic today. We need to be right there, ready to serve people wherever they are, with whatever the information is that they need in the form that they need it.

So our ability to change the way we generate the content and what form we publish it in needs to be really, really high touch. So this team at NetApp in particular is usually just slightly ahead of the curve. We’re very fortunate, we got a lot of really active and engaged people here.

And so the technology changes, the way we publish changes and the [00:08:00] way that we work changes as the needs of the business change. So now we have a group of people who are hyper-focused on agility, and accuracy findability. We have a differentiator, which is that we have the source of truth, right?

This is actually how you do something. In the environment that we’re operating in today that’s really extra super important that people know they can come and get the real straight thing from us, from our site created by people who’ve sat with the people who built the product and tested it. It’s a really important thing that we offer. So what we offer and the reason we offer it maybe haven’t changed that much. But the actual technical tools that we use in the processes that we follow. Have changed with the time. So we’re currently at a point where we’re really trying to focus on, like Adam I think alluded to this agility, right?

Speed, and the ability to keep pace with the business.

Justin Parisi: How does docs interact or work with the KB side of things? The knowledge bases? Does it interact at all?

Is there a synergy [00:09:00] there or is it just linking to each other?

Jen Kaufman: I mean, the teams work together really super closely. One of the things that we can see ’cause we can see a little bit of like, sort of who’s coming from where to our site, and quite a bit of traffic comes from KB to us and back and forth, right?

Mm-hmm. So people who are looking for information are using both sources of information to be successful. We’re gonna create content that’s ready at the release of the product, as we expect it to be used, the KB team is gonna cover all of the things that pop up after that in the course of using the product.

And sometimes we can even take some of that information and bring it back into the product docs and say, Hey, we learned something. We’re gonna add this to the documentation, because there’s a clearly a demand for it. And sometimes the knowledge base community will take content from docs in order to answer their questions.

So there’s a lot of synergy between the two teams. There’s a lot of, I think, appreciation between the two teams of what we do. We have a couple of different groups that NetApp, that create [00:10:00] content, technical content, and it’s all knitted together. You’ve got technical content being created by technical marketing engineers like somebody else.

You’re talking about me. Content being created by knowledge based experts. You’ve even got content on the community site that’s being generated by people who are deep into product usage. All of this works together, the hardware, universe, content.

These pieces all come together to give people the full picture and all the different sites and all the different channels deliver content for a different purpose, and together you get the whole picture of success for everyone who’s using the product. So, I don’t know if that really answers your question in a technical way, but I can just say that the synergy between the teams is extremely positive and one of the things that we try to do is make sure that we have provided good access to the content that we have through as many channels as possible.

Justin Parisi: And I was hoping it would lead you into your segue, the whole generative AI thing, like when we ask something, a question, I was wondering if, if I asked the knowledge base a question Yep. Can it take me over to the docs? If I asked the docs a [00:11:00] question and there’s a knowledge base article, can it take me there? So tell me about how gen AI is playing into this.

Jen Kaufman: Well if you want a really smart person to answer your questions about gen ai, that’s Grant Glass, but I can talk a little bit about the intent that our team had when we created our chat assistant, our AI assistant doc named after Doc Owl. Doc is our mascot. Doc is our wise owl who knows things. And Doc is trained on the content on docs.netapp.com.

So if you are chatting with Doc and you ask Doc a question about a product that we have provided information on. You can have a discussion with Doc. Doc is also trained on some additional content from other sources, but we add things to the dataset carefully and pointedly so that doc’s accuracy is maintained.

The point though was to make the information readily available, we’re trying to surface information wherever it is that people wanna find it. People want to use these [00:12:00] AI assistants. We want to provide one that’s gonna provide accurate answers. And so Mr. Glass here built one. So I think that might be a segue to talk to

Justin Parisi: I think it is. So Grant, how are we training AI to deal with our documentation? Because that’s a big job.

Grant Glass: Yeah. I think just the documentation outside of the kbs, we have north of 120,000 documents. And most of our products in our documents talk about very similar things.

We talk about storing and retrieving data. So to a large language model all of our products kind of sound the same. So we have to do a lot of work to say, well, what’s the value proposition of Keystone versus ONTAP versus Blue xp? A lot of times we have to push the model in certain directions based on what we’re seeing from the users. What kind of words the user’s using, if they’re using cloud [00:13:00] or on premises. We have a good indicator of what products and where they’re going. So, just disambiguating our own products and what our users and customers want is a monumental task, right outside of just even kb.

’cause we can think about KB is like special cases of, I have a particular type of implementation and even communicating a high availability pair in a standard cluster is gonna be pretty difficult task to do. And at the end of the day, we use the word ai, but there are large language models and the key there is language, right? This is why I think the writing team and the docs team is the best seated to have a generative AI solution. And we were the first public facing one for NetApp, because we deal with how we communicate about our products to the [00:14:00] user. And because we have that knowledge, we’re able to help the model navigate all of these documents and help users find the right ones to answer their questions. A lot of my data science background is essentially finding what are the clusters of words that surround unique products and how can we help models figure out from very sparse things from our users. Our users are very technical and they’re not very verbose people. So we’re trying to read a lot of our customer’s intentions and minds and that’s what’s the beautiful thing about general AI is now we have a direct line to see what are the kind of questions customers are asking. How are they interacting with our products, what are their pain points?

And it’s been a really interesting journey to have that interactive element through a chat bot and to Jennifer’s point about seeing what your users are [00:15:00] talking about and going back to the documentation to say, where can we clarify or where can we add more detail to help our customers better understand and use our products?

Justin Parisi: So what are you using to train the models? What sort of tools are you leveraging and, what does your workflow look like for that?

Grant Glass: The essential model that we’re using is called rag, which is retrieval augmented generation. So, it’s essentially trying to retrieve the right documents.

And working closely with Aksel’s team, we index and chunk and weight certain aspects to our documentation that we know indicate certain things. Is this a task-based document? Is this an overview page? What sort of topics is this document talking about? Retrieving the right document for the user is the key to the whole generative ai. I try to talk to other people about this is [00:16:00] like, you studied in college for many years and then it’s two years later and you’re sitting down for some calculus test, and maybe you kind of remembered calculus, but you kind of don’t. And so you try to answer questions, but you might get it wrong or you might get it right. That’s where a lot of people are saying, oh, well, generative AI doesn’t really know any answers or doesn’t get things right. What we’re doing in retrieval augmented generation is we’re giving you an open book. We’re saying, here’s a set of five documents.

We know the answer to your question exists somewhere in these documents. Find that answer. And only pay attention to these documents. And we’re confident in our retrieval. Aksel’s team’s done really good work on ensuring that we’re retrieving the right document. Once we have the right documents in play, it’s pretty easy to answer the question. The biggest learn that we’ve had [00:17:00] is, we think a lot about what’s in our documents. What are we trying to tell our customers? A lot of the times we’re not thinking about what we’re not telling our customers like, oh, you can’t do this with one of our products. You can’t delete a volume without changing an aggregate first or vice versa. That’s something we’re learning and continuing to try to improve is to try to say, well, we’re doing a really good job communicating, how can we essentially teach a model what you can’t do with NetApp products. What are the restrictions that we have? And that’s what we’ve been working on.

Justin Parisi: Is that kind of the fine tuning process there?

Grant Glass: Yeah. Yeah. It’s also understanding how customers use our language. They’ll say AFF or they’ll say 250, right? And what we’re doing by tuning doc is to say, oh, when you say 250 we know AFF A series 250 ONTAP hardware system. What we [00:18:00] do a lot in terms of training is, customers use a particular type of thing or they’ll say ASA. Great, we’ll assume that’s a ASA R2, because it’s the newest product.

So we do a lot of tuning based on user input, but also tuning to our documents. It’s kind of a holistic system of like, oh, we need to go back to the document. And we have a lot of users that are confused about workload versus workflows. So let’s go back to our documents and clarify those concepts a little further.

Adam Newton: Yeah, I wanted to supplement to what Grant was saying too. And this may be bridge toward Aksel’s team’s work, but the rag data set that Grant alludes to is something that we’re continuously updating. So it’s not a data set that’s set in amber and that we update every so often when we can, a monthly or whatever. It’s every day and it’s based on our builds. And our builds are happening every minute, every day. So the rag data set [00:19:00] is continuously updated and ready and agile for doc to use it to answer questions.

Justin Parisi: One thing that people are challenged by and that becomes a concern is the accuracy of the information that gets returned.

And I know I’ve run into this before where you ask an AI something and it comes back with something it thinks is really the answer, right? It’s very confident. And if you take it at face value, sometimes it’s not always right. So how accurate would you say your documentation models are today? And what sort of things are you doing to enhance those models? Do you have humans that are going back and checking? Is it customer feedback?

Grant Glass: I’ll take that question. We look at every single piece of feedback that comes to us, and we look at it as individuals. What you’re trying to figure out is did the generative AI element get this wrong?

Or, does this not exist in our documents? Or, if I read our documents, is this concept confusing? ’cause if it’s confusing to me, it’s a hundred percent going [00:20:00] to be confusing to generative AI bot. Or there could even be the way that the customer thinks about it. It’s obvious to me, but it’s not obvious to them.

And so that’s why it’s really important in this group. I think a lot of people think about the point of failure being the bot. But what you wanna do is you’re getting a good feedback to go back to the document and say, oh, we need to enhance the document. We know this from Adam’s experience of talking about the rag data set.

We had to do a lot of testing and we had to change titles to documents based on what we were seeing in ancillary generation because to someone that’s unfamiliar to NetApp, e-Series sounds a lot like ONTAP. So we do a lot of testing right along this, and this is where the human is absolutely necessary. This is a not an automated process. It’s figuring out where the point of failure is. We’re [00:21:00] not disambiguating the terms well enough in the AI solution. Are we not retrieving the right document? Do we need to tune our retrieval? Do we need to adjust the document?

So there are a lot of potential points of failures, but it’s also a holistic system. If you change a piece of documentation, we’re also gonna change the way that the bot operates based on what we’ve changed in the underlying documentation. It really requires a human being being present in all of those processes.

Jen Kaufman: And I was just gonna chime in on something Grant was touching on, which is the source content, in this case, the documentation can be optimized to perform well with one of these systems. And historically, let’s just say for the sake of argument, the content has been optimized for humans to read or for maybe search engines to give results on. Now we’ve added yet another audience [00:22:00] for the content, and that’s these large language models that need to be able to understand what we’ve written. And so there’s a certain amount of going back to the content, which in the case of ONTAP, there’s a huge amount of content that’s built up over the years, however long ONTAP has existed. And make sure that the content is being written in such a way that it’s optimally consumed by all of these audiences. So now, you’ve added yet another important consumer of the content that you have to optimize for. It’s important for the authoring organization to be kept current.

And luckily we’ve got great partners to help us do that. But it’s really important that a human do that synthesizing and then give instruction for going forward, how do we optimize this content for yet another new audience.

Justin Parisi: So Aksel, I know that your team is responsible for help helping to power this whole solution, right? Are your customers the docs team, or is it like NetApp customers?

Aksel Davis: Well, we actually are split a little bit, to be honest. But yes, of course our NetApp customers are first [00:23:00] and foremost. And we have a data science team that’s been formulated over the past two years to help coordinate a lot of those technical details to allow our engineering team to really focus on the infrastructure and the tooling to help support that. We are trying to build an AI landscape that supports different types of AI processes, because it’s not just about generative ai. Certainly that was our first ambition and we’ve accomplished that. And as Grant pointed out, we were the first at NetApp to do that on a public scale.

However we do have other areas we need to get into, and we’ve talked a little bit about supporting our writers, and giving them feedback and how they can improve their documentation. We also do that with engineering, when we get into our autogenerated content and every company in the world has auto-generated content, whether that be API, CLI, for those that are aware of Swagger ui, right? Engineers are responsible for generating a lot of that content. And engineers aren’t writers. We don’t always get it perfect. The AI system is supposed to help interpret and translate that into something that’s meaningful to a customer. We have [00:24:00] styles and guidelines and all of these other details that we factor in. We lent those engineering specs and we transformed that into our standard documentation that we published to docs.netapp.com and historically that’s been a very manual human process to go through and edit and review that. With ai we’re able to build types of systems that can create, edit, or standardize our engineering specifications and turn that into meaningful documentation. So that’s one critical area that we’ve been focusing a lot of our time on. We’ve been using Azure up to this point, and we’re looking at other technologies to supplement that. So for example, when the AI boom happened, everyone was focused on embeddings. That was the hot topic. Vectorization and embeddings. And we chose to go down a slightly different route where we did not use embeddings or vectorization. Now we’re looking at mainly because there were some costs and the frequency in which we republish and the expense of that now as it gets significantly cheaper and we have the time and energy to go [00:25:00] through and reevaluate. We’re able to supplement and we’re seeing multiple different data sources and transformation types that can be used in parallel based on the type of input you do. If you search for a command, the way you want that to return an answer is different than if you type a large question. The context matters of your query and the type of infrastructure that we use and the approaches that we use are different based on those inputs.

Adam Newton: One of the things that the data science team and development team have collaborated on and recently released in the last month for our writers is an authoring assistant that’s built in. It’s a plugin or extension of GitHub copilot. Aksel, do you wanna say a few words about that?

Aksel Davis: Yeah, absolutely. NetApp brought in GitHub copilot. By the way, we are centralized in GitHub. So as soon as we heard that, our ears perked up. We wanted to go check it out. And from a developer standpoint, you have all these different programming languages that it supports. And so you can imagine an analogous one for our authoring source, which is [00:26:00] Ask Doc. We basically wanted a way to create a plugin for it, that would enable us to run style checks and linting and other aspects on this content. And what was born out of that was a linting tool that we use that leverages GitHub copilot, that runs our style guide, that runs a few other checks.

And we look to build that out into our other linting application in the future, to build a comprehensive system that can do grammatical checks and other types of editing passes. The data that we collect is passed through back to our writers. We track that to make sure that we always have a human element that’s touching it before we would release any of that to docs.netapp.com. So that’s an exciting area. It uses slightly different technology with GitHub copilot than with our generative AI tool. So that’s yet another area of AI application that is slightly different to build out that ecosystem.

Justin Parisi: Were you saying linting, like a lint brush?

Aksel Davis: Yeah, so in engineering, we sometimes refer to something called a linter, which might be something that looks [00:27:00] at formatting. In coding, you might have certain indentation rules or if you’re doing Java, you might forget your semicolons or something like that. So anything like that, where it’s just running through trying to standardize your code base. And the same thing is true for our writing, where we go and clean up certain grammatical, stylistic changes.

Justin Parisi: So it’s a lint brush for code.

Aksel Davis: Exactly.

Adam Newton: And we are not rolling out the belly button snooper either, as part of that.

Justin Parisi: You could just use like scotch tape.

Adam Newton: Yeah.

Grant Glass: I think one of the things that Aksel’s touched on, and it’s very unique to this team, is one of the reasons that generative AI is actually pretty challenging in implementation is it requires both coding knowledge to be able to use the APIs and the environments to have some sort of call to these models.

But also, the prompting is a very large part [00:28:00] of getting the model to do what you need it to do, to get it to behave properly. And we’re fortunate enough to have writers that can be very specific and precise in their articulation of our standards. Like, what does a good lead paragraph have? And getting very specific, like lead statement needs to have four sentences, number one needs to have this or, we want a title and the title needs to have the proc name and an imperative verb. And we give the model, here’s examples of imperative verbs. Here’s an example of a good lead statement. The unique portion about Aksel’s team and Jennifer’s team and the data science team working together is we provide a solution that is robust in its code and it’s prompting in order to really create a good gen AI solution. That’s where I think I see other [00:29:00] organizations needing to open up a little more is involving your writing teams in the prompting because the writing teams are going to be able to compactly and succinctly tell the model what to do and what we’re looking for.

And so that’s, I think a huge learn that we had, especially from the Nadia, is how well writing teams and programming teams can work together to provide a really robust solution to these sort of writing problems.

Aksel Davis: Yeah, absolutely. And it’s not just that, but when we talk about the valuation techniques that you asked earlier, you might, for example, give question, answer pairs as part of your evaluation technique. And those writers can take customer questions or queries if you’re tracking that. And sometimes you wanna measure against that. You wanna know exactly what a customer’s putting in, they’re probably not gonna be always perfect, so that’s one type of test suite. Another type of test suite is having your writers go in and modify those questions to be more [00:30:00] accurate or relevant to what your documentation’s about. And that can actually be used as a mechanism to evaluate and train your model as well. Writers play a crucial role in that and they may not even realize it yet. So that might be something that other companies or people need to help empower their writing teams to do.

Justin Parisi: You mentioned a couple of technologies you’re using to help your team. What sort of net NetApp technologies are you leveraging for this work?

Aksel Davis: We already talked about some of the external tooling, like GitHub copilot and open ai. Within NetApp we’ve done some experiments with certain products, some that have been released, some that haven’t been released. And I would say that in all of our evaluations, each one has its own merits and benefits.

But because we had already been at the forefront of how we were doing it we were really looking at what are the net value adds. So things like classification was really important to us. That was not an area that we had historically been at the forefront of. Segmentation was really important to us, how we segment our documentation and we felt that our strategies were [00:31:00] comparable to the end product solutions that were there.

And so it lended pretty naturally into being able to tie in and evaluate against those systems. So I’m not sure we’re in a position to chat too much about that. But I will say one thing is that we have a partnership with each of the different first party cloud providers. So we are actually working with AWS right now to ingest their documentation from AWS through an MCP server and convert and ingest that into our doc rag data set. So very soon we will have AWS documentation flowing into doc and into our site search. So we’re very excited about that. And that’s one area that we’re improving our tooling.

One big exciting project we did over the past year was we’ve been really revamping our build infrastructure. As Grant alluded to, we have over 120,000 documents in English, and then we have that translated into nine languages. So we’re publishing over a million documents, and we will refresh much of that content on a daily or even weekly basis. To do that, we have to have a lot of build [00:32:00] runners. It takes a lot of time to be able to transform that. We transform it to PDF as a secondary output format. And with all of that combined and then indexing into search, builds originally could take about an hour and a half, two hours, sometimes even three hours like our ONTAP build, for example. To cut that down, we went to an incremental publishing approach and we were using tools like Trident and Kubernetes and Ansible to be able to help speed up our build infrastructure, but also in our publishing to get down to an incremental build.

We use GitHub actions for that. It’s been an absolute amazing tool to help speed that up. We now have, I think 30 build runners, and if we have to do a full site rebuild we can spin up a hundred plus build runners which would normally take five days.

To build all of our sites, we can now do an about an hour or two. We can either do that in two different environments, one being Kubernetes based, one being an Ansible based. And we have both of those available to us as needed, an ad hoc basis. We put in Trident and we found some great additional benefits to that as well, which is a NetApp [00:33:00] product for containerization. Those are some really key areas. And then looking at Amazon fSX ONTAP, which we use for a lot of our storage for our publishing environment and hosting.

Justin Parisi: So you’re running in this all on NetApp. I just wanted to make that clear. You can do all this on NetApp. I mean, if you chose. So one of the challenging things about doing anything with ai, whether it’s developing AI or using AI internally to improve your documentation, is navigating the intellectual property, navigating the legal challenges that might come into play.

’cause nobody wants to have their internal knowledge out in the public domain. Like they want to try to protect that as much as possible. Adam, how are we navigating that landscape at NetApp? What sort of things do we do to try to protect that ip?

Adam Newton: That’s a very good question Justin.

At the outset when we were working on our generative AI chatbot, other teams at NetApp were looking at gen AI too, and [00:34:00] wondering how they might leverage it or build solutions from it.

And so we engaged with those teams across functional work group that included our legal department and talked to them about what would be appropriate use. The advantage we had, the we being the docs team, was several years before as part of that open mindset, published all of our content out on the public internet. So what we were publishing was not private in any way or proprietary. It was out there in the wild, so to speak. So that was a major hurdle. People listening to this may be in different circumstances, they may have different policies and governance models at their companies, but for us it was a fairly straightforward move to take publicly available content and basically recombine it through the magic of Gen ai. Now, I will say, on the user experience front though we did design into our ui some appropriate caveats basically about what you [00:35:00] are interacting with as an end user. We do try to provide guardrails and guidance and make sure it’s clear to everybody that they should always verify what they’re seeing in terms of the answers. And I think Grant alluded to it, and our generative AI solution doc, the chat bot, all of its answers are buttressed by citations. We encourage people to always verify.

Justin Parisi: So I shouldn’t just blindly format my cluster because the gen AI told me to do that? That’ll fix your problem.

Grant Glass: I would just say the last thing. We have a really great partnership with Microsoft and one of the reasons why we were able to do this solution was we were working on it for two years in open AI in Azure and that provided enterprise level security and privacy that we needed. The first version of this used ANF to store all these PDFs and to see before we hook this into our publishing environment, is this [00:36:00] solution gonna work for our type of documentation? Having a service, like ANF and an open AI instance provides some rapid prototyping. So, I think we are also at an advantage because of that close partnership we have with Azure that we are able to start early. What you’re seeing now is two years of learns from those smaller prototypes to the customer facing enterprise level app that we’ve deployed now. So I think that was a huge part of that initial journey that we are doing and we are enabled by that partnership we have with Microsoft.

Adam Newton: Yeah, that was a good call out, Grant, and you mentioned ANF, Azure NetApp files, that was part of it. Maybe one thing I’d say, I alluded to this, mindset and culture change for our broader team, going from a walled garden specialty team or artisanal sort to a more open source and open mindset.

And I’d say another mindset change for us, [00:37:00] because of Gen ai and because of the work we’ve done on docs is that, and Jennifer alluded to it earlier as well. Generally speaking, we see our roles not just as content creators and publishers, but as data stewards. That’s something that we have been talking about internal to our team. What does it look like in this new gen ai future to own a dataset, govern it, maintain its currency and consistency, continually groom it and architect it in such a way that it’s ready and versatile for other solutions that are not yet know fully visible on the horizon. So I think that’s a very interesting change for us and having some data scientists on our teams and having an engineering team, people whose minds tend to be more technical, who think in terms of systems and governance has been extremely helpful for our team to be able to make that change.

As part of that openness that we embraced, we have invited people like tme into our system. One of the benefits that customers have now is finding content in one place and then having [00:38:00] having the content that was previously sitting on people’s hard drives in word files and whatnot, and then being published occasionally as PDFs is now part of an entire, ever continuously being built dataset, right? Content that was previously segregated and out somewhere in a different website is now just part of this big data set that we can continuously build solutions on top of.

Justin Parisi: All right, so what sort of new things you might be coming out with, do you have anything to talk about there?

Adam Newton: Nah, we’re done Justin. We don’t have anything else to do.

Justin Parisi: That’s good. That’s good. I’m glad.

Adam Newton: Everybody got an A, and we’re all taking vacation now.

Justin Parisi: Mission accomplished.

Adam Newton: Yeah, all joking aside, I think one interesting aspect that we have ahead of us and the opportunity for the company and for our customers would be to join our dataset to other data sets.

So, for example, there’s a very large data set that comes back from our auto support feature. And I think how might we leverage that data set in ours in perhaps an agentic framework is something [00:39:00] that holds promise in the future.

Justin Parisi: Something like using auto support to back up a best practice recommendation, saying, sure, 95% of our customers use this best practice.

Adam Newton: We see what’s happening in your environment. Most people in your particular situation, most people take this path and here’s how to do it.

Justin Parisi: Yeah.

Adam Newton: Something along those lines.

Justin Parisi: Informed information gathering there. What about tying into something like system manager, Blue xp, whatever we’re calling it these days, having recommendations pulled directly in from Asup.

Adam Newton: Yeah, we roll up or report into the chief design office here at NetApp. So our partners are UX researchers and UX designers, and that’s been an incredibly fruitful partnership opportunity for us where now we talk to our designers, we work in Figma with them, and imagine new experiences in products.

So yes we’re imagining some kind of assistant inside of our products, that our content is [00:40:00] essentially syndicated into. And combined with again other data sources. So yeah, you’re barking up the right tree and that’s absolutely something our Chief Design officer Christina Storm is passionate about.

And that Aksel’s team, from a technical perspective very, very well positioned to enable.

Justin Parisi: We have a team that mines asop and sends out emails and alerts for things that aren’t being followed. Best practices, so that if you have a volume that maybe is like getting to be full mm-hmm.

Hey, by the way, your volume’s getting full. Here’s the kb, or here’s the documentation telling you about why that’s not great and here’s how to fix that. I think that kind of informed use case might be a good way to tie all of it together.

Adam Newton: Yeah. And Justin, I also wanted to say, we haven’t focused at all on it in this call, and it’s not the central purpose ’cause we’re talking about docs.

But the other team that reports to me, the globalization, I wanted to point out are also using ai specifically to evaluate the quality [00:41:00] of the machine translations. At NetApp, 94% of all the content that’s delivered in language is actually just machine translation. So 6% is human done. Our goal in this year is to get to a hundred percent. But I think it’s very interesting what Edith Bendermacher and her team and the GPSO team, the globalization team at NetApp are doing to essentially evaluate quality of machine translation and continuously improve it through AI evaluation.

Justin Parisi: Yeah, that’s one of those mundane tasks that I can imagine that having an army of humans do is just mind numbing.

Adam Newton: You’re right. And the content reviewers in language, Brazilian, Portuguese or Japanese or whatever will still always be part of that loop.

But the idea is that it’s a machine first and AI led translation workflow as opposed to human.

Jen Kaufman: One of the other things that we were talking about, content offered up in product, and I was gonna say one of the things that we’re also working on and planning for is to [00:42:00] incorporate search and AI functionality on the docs pages so that you get the benefit of both AI summarization and the opportunity to chat with the assistant and also search results combined and offered up together on our site, on docs.net.com. And that is also something that potentially could be hosted from within a product as well. So you get the benefit of all these different engines working together from within the product.

So that that’s very future think. But it’s something that we’ve talked about and something that we’re starting to look at implementing to some degree, at least on the doc site to begin with.

Adam Newton: That’s a good one. There are some interesting subroutines too, where we’re generating content from engineering specs, how might we through ai create synopses of what’s happening reports about deltas but articulated them in narratives and giving them to business stakeholders.

So connecting our linting to not just the technical aspects of what it does, but to then [00:43:00] build business intelligence. Here’s what we see as we’ve been processing all of this content. One last thing. I think Aksel alluded to it. We have first partner agreements with Amazon, Google, and Microsoft. And I think MCP servers seem to be a thing we are already leveraging for integrating the Amazon FSxN docs into our rag dataset. I think we have an opportunity to help our customers and our internal stakeholders interact with their content in our context through MCP and other syndicating technologies like that.

Grant Glass: It’s finding ways to have all of these disparate data sources talk to one another and one way around that is MCP servers. That’s definitely the next sort of thing. How do you integrate even more data sources, because we understood that it was pretty difficult to do. Adding in something like KBs is even a more [00:44:00] monumental task. So having these MCP protocols, I think helps organizations implement all of these disparate data sources together in one agent flow.

Justin Parisi: Wasn’t MCP the thing in Tron, isn’t it? Am I misremembering that?

Adam Newton: I don’t Anyone, anyone.

Justin Parisi: I feel like it was

Grant Glass: yes. Master control program. Yeah. Yeah. It was the big cylinder, right?

Justin Parisi: I don’t like that.

No. Me people that for you. No. This is, we see where this is going,

Jen Kaufman: But I mean, what we’re trying to do all across this organization is apply AI intelligently and you’ve heard us talk today about applying AI to help people find and use the content. Applying AI to generate and produce content to speed and make more excellent. We’re just trying to apply it in places where it makes sense so that we’re speeding our business and making the experience of the customer more excellent and not just AI everywhere, which is [00:45:00] not very helpful, right? But it takes thought and it takes practice and it takes experimentation and there has to be an actual use case for it, and then we can put our energy towards it. There’s a lot of activity and we don’t always know what the future will bring. But I think as a team trying to collectively experiment with whatever we think shows promise to see if there’s anything that we can extract from it. I know, Aksel, is there anything that you wanna touch upon that your team is talking about, maybe even in the agentic space or anything else?

Aksel Davis: The biggest goal for us is about findability of our docs. Justin talked about that on the onset. We’re coming into an age where GDPR and privacy is so important to our customers and to the global population. It becomes harder and harder for us to trace who is using our documentation.

It’s the number one question we get asked by our senior executive leadership, who is using your docs, who’s coming there? And we have to sheepishly say. We don’t know, right? Because we can’t track that information at a certain level. So we’re hopeful that we can at least see trends and patterns and how people use our documentation, and that can [00:46:00] help inform some of the decisions that we make and some where we place investment into our technology stack going forward. I think that’s the biggest opportunity we have and how we partner with other departments because everyone wants to build their own AI system. Everyone wants to quote, own the content ecosystem, right? How do we partner together across the NetApp eco enterprise to build a centralized solution for our customers?

Because ultimately that’s what we’re after. Making it as findable as possible. And so that’s our biggest charter over the next year and I think we’re well equipped to go there.

Justin Parisi: Yeah, I think you guys are on the right track there with this. All right, so NetApp Docs team, thanks again for joining me and talking to us all about what’s going on with Gen AI and the NetApp docs as a whole. Again, if you wanted to reach any of these fine folks, they’ve given you their contact information at the beginning, and we’ll include information with the associated blogs or LinkedIn post or whatever we have here, so you can contact us that way. I wanna, again, thank Adam, Jennifer, grant and Aksel for joining us today. And thanks for listening to the [00:47:00] techontappodcast.com.

We will be happy to hear your thoughts

Leave a reply

Daily Deals
Logo
Compare items
  • Total (0)
Compare
0
Shopping cart