Finally, an Open Standard for the Karpathy LLM Wiki is HERE — Transcript
Full transcript
- 0:00A couple months ago, Andre Karpathy
- 0:02released the idea of the LLM wiki. It's
- 0:04a pattern for building personal
- 0:06knowledge bases using LLMs and it
- 0:09totally took off and for good reason.
- 0:11There's a lot of power in the simplicity
- 0:13here. So, this single markdown document
- 0:16in GitHub called it gist got to 40,000
- 0:18stars. And seriously, you can take this
- 0:20file, copy it, paste it into your coding
- 0:23agent, and ask it to build you an LLM
- 0:25wiki and it's going to be able to just
- 0:26basically oneshot it. So, it's really
- 0:28easy to get started. And the idea here
- 0:31is when we're building a personal
- 0:33knowledge base for our second brain,
- 0:35instead of just dumping in a bunch of
- 0:36documents or indexing things for rag, we
- 0:39can have the LLM help us build something
- 0:41smarter, incrementally building and
- 0:43maintaining a persistent wiki with
- 0:45structured interlink collections of
- 0:47markdown files. And so the idea here is
- 0:50as we're adding in more sources over
- 0:52time like meeting transcripts, plan
- 0:54documents, articles from online, it's
- 0:56going to not just index it, but it's
- 0:58going to read each file, extract key
- 1:01information, and integrate it into the
- 1:03existing wiki. So updating things like
- 1:04the entity pages that it creates over
- 1:07time, so we have that knowledge graph
- 1:09for agent to traverse through and
- 1:11remember all the important information
- 1:12that we're bringing in. So, at this
- 1:14point, pretty much everybody is building
- 1:16their own LLM wiki in their second
- 1:18brain. But this isn't enough. And the
- 1:21main problem that we have here is when
- 1:23you take this gist and you build your
- 1:25own version of an LLM wiki, it's going
- 1:27to be structured differently than the
- 1:29next person doing the same thing.
- 1:31There's no standard. And so, there's
- 1:32really not a way to share your LLM wiki
- 1:35with someone else. And that's a bummer.
- 1:37You can think of a lot of different use
- 1:38cases where you'd want to curate a
- 1:40knowledge base over time and then share
- 1:42it with other people like other people
- 1:43on your team. Maybe you want one wiki
- 1:45for the team that everyone's second
- 1:47brains are accessing independently.
- 1:49Maybe I want to create a wiki for my
- 1:51YouTube content and then share that with
- 1:53you. There are a million reasons. But if
- 1:56your agent doesn't know exactly how I've
- 1:58structured my wiki with the different
- 2:00metadata and my entity files, it's not
- 2:02going to be able to search through it
- 2:03optimally. We need a standard so that
- 2:05everyone's building wikis in the same
- 2:07way so that we can share them freely.
- 2:10And so that is what Google has released
- 2:12here with their open knowledge format.
- 2:14It is a beautifully simple thing just
- 2:16like Harpathy's LLM wiki idea where it's
- 2:19just a simple standard built on top so
- 2:22that you can guarantee you're building
- 2:24your wiki in a way where other people's
- 2:26second brains can understand it and vice
- 2:29versa. And so in this video I want to
- 2:31cover why OKF is so powerful. It really
- 2:33is the future of personal agents. And I
- 2:36want to show you how easy it is to get
- 2:37started with this standard, both for new
- 2:40LM wikis and even transferring existing
- 2:43ones into this format. Very easy to do
- 2:45that. And no matter the wiki, no matter
- 2:47how much you're going to share it or
- 2:48not, this is important even as an
- 2:50optimization on top of Karpathy's LM
- 2:53wiki idea. And I know that Google is
- 2:56lagging in the AI race right now. Gemini
- 2:58is not as good as GPT and Claude, but
- 3:01they have been releasing some really
- 3:02good stuff on how to leverage LLMs
- 3:04effectively. And I think that's a
- 3:06totally different lane than building
- 3:08LLMs. Well, so I think this is something
- 3:10really worth leaning into even if OKF
- 3:13doesn't end up becoming the standard
- 3:14down the line for personal agents.
- 3:16There's going to be something like this.
- 3:18And so it's good to understand this now.
- 3:20Okay. Now, let's really get into OKF. So
- 3:23there are two things that they're
- 3:24standardizing here. The first is how we
- 3:27are organizing information like our
- 3:28entity documents and our concepts. And
- 3:31then the second standardization is the
- 3:33exact fields that we're going to have in
- 3:35our metadata. So this is the information
- 3:38that we tag at the top of every single
- 3:40document to give the agent a richer set
- 3:43of information. So we can even like
- 3:44query based on the title or the tags. So
- 3:47we have categorization. This is one of
- 3:49the most important things to let the
- 3:50agent traverse through our wiki like a
- 3:52knowledge graph. And really the best way
- 3:55to make this concrete for you is to show
- 3:57you what a traditional Karpathy wiki
- 4:00looks like. So we'll take a look at
- 4:02this. This is one of the first wiks that
- 4:03I built when Karpathy released this
- 4:05idea. And then we'll get into some of
- 4:07the problems that we have here. So at
- 4:09the top of every single wiki is your
- 4:12index file. You have the agent maintain
- 4:14this every single time it's bringing new
- 4:16information in. And the index file, it
- 4:18reads this when it's first searching
- 4:20through your knowledge base, pretty much
- 4:21every single time. And so this just
- 4:23gives you a high-level overview of all
- 4:25the documents that you have access to in
- 4:28the wiki. So the article and then a
- 4:30quick summary so it knows if this is
- 4:31something that it should look into based
- 4:33on the user's request. And so every
- 4:35single time we add in new documents,
- 4:37this is evolving. And so the agent will
- 4:40read this and then based on what we
- 4:41asked it to do or the question, if it
- 4:43figures like I should look at superbase
- 4:45o this concept right here, this entity
- 4:47document, then it'll drill into this. We
- 4:50also have the metadata like I talked
- 4:51about earlier like the title and the
- 4:53tags so that it can also search based on
- 4:56this like if it wants to look at the
- 4:58category of security then it can filter
- 5:00out just those documents and so then we
- 5:03have the full sort of like skill.md here
- 5:05this is like progressive disclosure like
- 5:07skills where the index tells it the
- 5:09knowledge it has and then it can read
- 5:11the full document if it's appropriate
- 5:13and then we also link to related
- 5:15concepts down here and that link is what
- 5:17really gives us this graph view where
- 5:19You can see how all of our entities and
- 5:22other documents are connected together.
- 5:24So, the agent can sift through this to
- 5:26really get a comprehensive set of
- 5:27information if the question really calls
- 5:30for it. And so, looking at one of these
- 5:32documents here, it might feel like it's
- 5:35overwhelming to build up all this
- 5:36knowledge over time, but seriously, with
- 5:38an LLM wiki, you are just giving the
- 5:40reins completely over to an LLM. So, you
- 5:42don't have to be technical. You don't
- 5:44have to spend a lot of time maintaining
- 5:46this. Literally the whole benefit of the
- 5:48wiki is that up until we've had LLMs for
- 5:52this, it was way too tedious to create
- 5:54this sort of knowledge base where we're
- 5:56responsible for understanding related
- 5:58concepts and building that over time as
- 6:00we're adding in new information. Like
- 6:01there's so much tedious work here that
- 6:03LLM are really, really good at. But as
- 6:05much as they're good at this, they
- 6:07aren't going to create this system in
- 6:10the same way that someone else will with
- 6:12with their LLM, right? like the way that
- 6:14we link related concepts might be
- 6:16different. The way we structure
- 6:17information, even the metadata, like
- 6:20what if we don't have tags, but we have
- 6:22a field called categories. I mean, even
- 6:24something as simple as that, that that
- 6:26little change might make it so that if I
- 6:28gave the knowledge base to another
- 6:29person's agent, it wouldn't know how to
- 6:32search through things categorically. It
- 6:34would have to dive into the metadata
- 6:36first to understand that, and it might
- 6:38not decide to do so. I mean, all these
- 6:40little problems will start to compound
- 6:42when you don't have the same metadata,
- 6:43you don't have the same folders. That's
- 6:45what we're looking to do here with OKF.
- 6:47All right. So now, if you want to build
- 6:49with OKF, create a new knowledge base
- 6:51with this format or even refactor one to
- 6:54use the open knowledge format, look no
- 6:56further than their spec.md file. So,
- 6:58this is in their repo. I'll link to it
- 7:00in the description. This is just like
- 7:02Karpathy's gist where you copy this
- 7:05document. Like you literally just click
- 7:06this one button right here, put it into
- 7:09your coding agent, and tell it to either
- 7:11build you a wiki following the open
- 7:12knowledge format or even refactor an
- 7:14existing one. Like I said, it's going to
- 7:15knock either of those out of the park
- 7:18because this is kind of like a skill. It
- 7:20teaches the coding agent everything it
- 7:22needs to know about the standard. Like
- 7:24here is the terminology. Here's how we
- 7:26structure the bundles. I'll show you
- 7:27more on this in a little bit. Here is
- 7:29how we build the YAML front matter.
- 7:31different attributes that we have for
- 7:33each one of our documents like the tags
- 7:35for categorization, right? Like this
- 7:37single source of truth is all that it
- 7:40needs. And because it's such a simple
- 7:42format, a simple standard overall, it's
- 7:44not really going to get confused going
- 7:46through this. I mean, it's a pretty long
- 7:48file, but in terms of what large
- 7:49language models can handle these days,
- 7:51especially with, you know, GPT 5.5 or
- 7:53Opus 4.8, this is not much instruction.
- 7:56And it it also doesn't really matter the
- 7:59scale of your current knowledge base if
- 8:01you are refactoring because you can
- 8:03specifically ask it to use sub agents to
- 8:06work through the different sections of
- 8:07your knowledge base to refactor it to
- 8:10this format. So really easy to scale,
- 8:12really easy to just have the agent rip
- 8:13through this spec. The sponsor of
- 8:16today's video is Post Hog, a single
- 8:18place for you to understand how users
- 8:20are actually using your application to
- 8:22debug and fix issues and test and roll
- 8:24out all of your changes. And I'm excited
- 8:26for this because I am using Post Hog
- 8:29myself in Archon, my open-source AI
- 8:31coding harness builder. I'm legitimately
- 8:34leaning on the data insights that I get
- 8:36from Post Hog every single day so that I
- 8:38know exactly how to improve Archon in
- 8:40the way that users actually need. And
- 8:43installing Posthog is incredibly easy.
- 8:45You just click on the install with AI
- 8:46button on their homepage that I'll have
- 8:48linked to in the description and boom,
- 8:50it's a single command you can run a
- 8:52wizard that will essentially be a senior
- 8:54engineer helping you set up analytics
- 8:56for your entire application in just
- 8:59minutes. And you can also create custom
- 9:01data views like this is the dashboard
- 9:03that I'm looking at every single day to
- 9:05see how people are actually using
- 9:06Archon. And then we can also drill down
- 9:08to get very granular as well. So the
- 9:10individual runs of Archon, I can click
- 9:13into this here to see all the details.
- 9:15And so we can go very high level all the
- 9:17way to individual parameters as we need.
- 9:20It's got the analytics for everything.
- 9:22And so production is the time where you
- 9:25can't be flying blind. When you have
- 9:27something deployed out to the world, you
- 9:29need observability. And Post Hog is the
- 9:31best for that. So I'll have a link in
- 9:32the description. I would highly
- 9:34recommend checking them out. And I
- 9:36talked about this a little bit at the
- 9:37start of the video, but this really is
- 9:39the future of personal agents. It's like
- 9:41what MCP did for agentto tool
- 9:44communication, this OKF is doing for
- 9:47agent to knowledgebased communication.
- 9:50And one of the most important things in
- 9:51the spec here is that they talk about it
- 9:53being a standard both for consuming
- 9:56knowledge bases like searching through
- 9:57them, but also producing knowledge
- 10:00bases. How do we evolve the wiki over
- 10:01time? build up the entity pages like
- 10:04Karpathy talked about in the initial
- 10:06gist. We really are building on top of
- 10:08it. And one of the really interesting
- 10:10things to think about here is yes, this
- 10:13is fantastic for sharing knowledge bases
- 10:15or having a teamwide knowledge base.
- 10:18This is also really good though even if
- 10:19you're never going to share a knowledge
- 10:20base. Think about this. If everybody has
- 10:23the same standard for how they are
- 10:25building up their own personal knowledge
- 10:27base, everyone can share ideas more
- 10:30like, oh, here are the entity pages that
- 10:33are working really well for me and this
- 10:34is how I want to organize things under
- 10:36the standard. And then because you have
- 10:37the standard as the foundation, it's
- 10:39easier for other people to take those
- 10:41ideas. And so what we're also I think
- 10:43what we're going to see is like yes, I
- 10:45don't think OKF is going to in the end
- 10:47be the standard, but we're going to see
- 10:48something like that and we're going to
- 10:50see the standard evolve over time so
- 10:52that it's easier and easier for people
- 10:54to create these really rich knowledge
- 10:56bases without having to spend a lot of
- 10:58time upfront designing it with the LLM.
- 11:01Now, of course, sharing wikis with other
- 11:03people is the biggest benefit of OKF.
- 11:05And that leads me into the example that
- 11:08I have for you that's also a gift I'm
- 11:10very excited to share. I have built a
- 11:13bundle, that's what you call an OKF
- 11:14Wiki, that packages up all of my
- 11:17favorite AI coding YouTube videos on my
- 11:20channel. And so, here's the thing. I'm
- 11:23excited for this. I know that a lot of
- 11:25you, you don't watch my entire video
- 11:27every single time. You're going to sift
- 11:29through things. You're going to just
- 11:30take the transcript and feed it into
- 11:32your second brain and ask questions. You
- 11:34guys are already doing something like
- 11:35this, but now making it easier for you
- 11:37because I'm prepackaging up sets of
- 11:40videos. I actually want to start doing
- 11:41this so that you can very easily bring
- 11:43it into your second brain and ask
- 11:45questions as it relates to what you
- 11:46actually care about or what you are
- 11:48working on specifically. And so take a
- 11:51look at this. All you have to do is
- 11:53first of all take this spec and give it
- 11:55to your coding agent. You have it teach
- 11:57itself OKF. And then you go to this repo
- 12:00with my AI coding knowledge bundle. I'll
- 12:02have this linked in the description as
- 12:04well. And you just paste this prompt
- 12:06into your coding agent. That's it. You
- 12:08give it the link to this repo. You tell
- 12:10it to read the readme and set up
- 12:12everything and it already understands
- 12:14OKF. So, it links those two things
- 12:15together. Brings the bundle into your
- 12:18local Obsidian or Notion or whatever
- 12:20you're managing your knowledge. And then
- 12:21boom, you can instantly start asking
- 12:23questions. You don't have to bring in
- 12:24the transcripts yourself. This is the
- 12:27easiest way for just content creators in
- 12:29general to share their knowledge with
- 12:31the world. They can create bundles. I'm
- 12:33creating bundles for all my videos now.
- 12:35And so this is just one example of what
- 12:37OKF unlocks for us. And so I'll also
- 12:40show you what this bundle looks like
- 12:41because it's a really good example of
- 12:43what OKF is really doing for us. All
- 12:45right, let's get into the belly of the
- 12:47beast. Now I'll show you how I've been
- 12:48setting up OKF and we'll get into the
- 12:50example bundle as well. And so something
- 12:52that I do for my second brain, every
- 12:55single system that I build in, I always
- 12:57have a tople document that talks about
- 12:59how it works. Like this is how I'm
- 13:01working with OKF bundles. And then here
- 13:03are the different bundles that I have.
- 13:05So I basically have an index so it knows
- 13:07the different bundles that it can go
- 13:08into and search and read the index that
- 13:11we have in there. So we kind of have
- 13:12like two layers of indexing. And then I
- 13:15also built a simple CLI script. This is
- 13:18actually there in the example bundle
- 13:20that you can clone that makes it easy
- 13:22for it to in the command line list out
- 13:24my bundles to view a specific index and
- 13:27then you know once it finds one of those
- 13:28files it wants to read then we have the
- 13:30command line tool to read by a specific
- 13:33bundle and concept ID. So I've added
- 13:36like a little bit of organization on top
- 13:38of OKF with just how I manage many
- 13:41different bundles but otherwise I'm
- 13:43following the format exactly. And so
- 13:45let's actually look at one of these.
- 13:46I'll click into bundles here and we'll
- 13:48go into the one that I just shared the
- 13:50GitHub for. So, if we look at the index
- 13:53here, we can see that I have two
- 13:55different sections and this is actually
- 13:57a smaller bundle. So, I didn't want to
- 13:59do something super complicated. So,
- 14:00there really are just two sections. I
- 14:02have the videos that I've put in this
- 14:04bundle, which it's it's rather small.
- 14:06There's only four videos, but these are
- 14:08like the best and most up-to-date ones
- 14:10on my channel for AI coding. And then I
- 14:12have the concepts as well. So different
- 14:14things that I talk about throughout
- 14:16multiple of the videos that I want to
- 14:18extract into its own entity page. And so
- 14:22the index here says here are the
- 14:23sections. And then I don't actually have
- 14:25a list of each one of the individual
- 14:27files because I'm just going to have the
- 14:29agent read the files that we have in
- 14:32concepts or videos, right? Like it can
- 14:33list out here are all the files or it
- 14:35can read the index within concepts
- 14:38itself, right? So, however you want it
- 14:40to navigate, it's going to be able to go
- 14:42through these different layers of
- 14:43documents or just do a keyword search.
- 14:45And so, clicking into any one of these,
- 14:47like the PIV loop, for example, this is
- 14:49the primary mental model that I always
- 14:51teach for AI coding. Very important to
- 14:53have a process for yourself to plan,
- 14:56implement, and validate whatever you're
- 14:58creating with a coding agent. And so, we
- 15:00have the YAML front matter at the top.
- 15:02And the type, this is what is required
- 15:04by OKF. It is the single required field
- 15:07in the metadata because this is what
- 15:09gives categorization to your documents.
- 15:12So like this is the type of concept. If
- 15:14I go to a video here, the type is video.
- 15:17So we can search over just the videos
- 15:18over just the concepts which is
- 15:20especially powerful once you get bundles
- 15:22that are a lot bigger than this. Again,
- 15:24this is just an example here. But then
- 15:26we also have all of the optional titles
- 15:28in OKF. So title, tags, related videos.
- 15:31This is how we link things together,
- 15:33right? Like you saw with that other wiki
- 15:34I showed earlier, it was just things
- 15:36were linked at the bottom. However, this
- 15:39now makes it so it's easier to navigate,
- 15:41creating a standard for how we are
- 15:43linking our entities together. And so
- 15:46each one of these are optional. Only
- 15:48type is required in OKF. But just
- 15:51because you don't always have these
- 15:53doesn't mean that your agent won't
- 15:54understand it, right? Like if your agent
- 15:56is a consumer of OKF, if you gave it the
- 15:58spec and taught it to be a consumer,
- 16:00it's going to know how to leverage these
- 16:02fields for better searching and
- 16:04traversing through the knowledge graph
- 16:06that we have here. And so then this is
- 16:08just all of our information on the piv
- 16:10loop. I kept it nice and simple. And
- 16:11then also linking to videos as well,
- 16:13which maybe is like a little bit
- 16:14redundant with related videos. So I
- 16:17could probably make this bundle a bit
- 16:18better, but I just wanted to have this
- 16:20as an initial example. And it is
- 16:21something that you can immediately bring
- 16:23into your second brain. Just start
- 16:24asking questions. Like I'll show you an
- 16:26example here in my terminal. So first of
- 16:29all, at the top level of my second
- 16:31brain, I just asked what bundles do I
- 16:32have? It ran a command here. So it used
- 16:35that little CLI tool to list out all the
- 16:37bundles that I have. And then it told me
- 16:39that and then I just asked it a
- 16:40question. So not even telling it what
- 16:42bundle specifically to look through. I
- 16:44said, "What's Cole's single biggest idea
- 16:46for getting reliable code out of an AI
- 16:48coding assistant?" and it ran four
- 16:50commands in total. So first of all it
- 16:52decided to read the coal AI coding index
- 16:55that's the GitHub that I have for you
- 16:57and then based on the index it knew like
- 16:59okay let's take a look at the concepts
- 17:01here and then from the concepts it's
- 17:03like okay the single most important
- 17:04thing I don't know what in the index
- 17:06told it that but it's like context
- 17:08engineering let's read the concept of
- 17:10context engineering so we can see the
- 17:12progressive disclosure as the agent is
- 17:14figuring out where it needs to look down
- 17:16to find the answer for me and then we
- 17:19get the final answer here So just
- 17:21beautiful to watch it work. When we have
- 17:23something structured like this, it's so
- 17:25easy for it to start with really not
- 17:27much context at all and then drill down
- 17:29into exactly what we need. That's what
- 17:31OKF gives us as a standard. All right.
- 17:34So if you're not sold on the idea of
- 17:36having a standard for the LM wiki at
- 17:39this point, I don't know what to tell
- 17:40you. The one critique that I think is
- 17:43actually pretty valid with OKF is a lot
- 17:46of people are saying that it's too
- 17:47simple, right? like there's not a lot of
- 17:49value or substance that's actually added
- 17:51on top of the Karpathy wiki. So, I've
- 17:54I've seen that a few times just as I've
- 17:56been doing a lot of research. I mean, I
- 17:57put a lot of time into prepping for
- 17:59these videos. I think it's kind of valid
- 18:00because if we look at like what it's
- 18:02really doing on top of the Carpathy
- 18:04wiki, it's it's speaking to like exactly
- 18:06how you organize your different files.
- 18:09Like they they specifically have like
- 18:10indexes within the folders and a top
- 18:13level index like you saw in my bundle. I
- 18:15mean, that's something I didn't really
- 18:16have in wikis before. And then we have
- 18:18the specific fields in our metadata like
- 18:20the type is required. The other ones are
- 18:22optional but these are the ones that
- 18:24they recommend. Like that's pretty much
- 18:25it. It's how we organize and what is the
- 18:28metadata. That's pretty much all that we
- 18:30actually have in the standard. And so
- 18:34like the argument is kind of valid where
- 18:35it's like what is it really giving? Like
- 18:37there's there's not much there. But I
- 18:39think that's also the point, right? Like
- 18:41minimally opinionated. It's the bare
- 18:44minimum layer that we need on top so
- 18:47that we can produce and consume these
- 18:50wiks in exactly the same way across
- 18:52everyone's agents that lean into OKF.
- 18:55Like I think that's actually a good
- 18:56thing. I think that's a benefit, not a
- 18:58downside. The fact that there's not much
- 18:59substance here might seem
- 19:01counterintuitive, but I think that is
- 19:03actually a good thing. And I encourage
- 19:05you just try out the bundle that I have
- 19:08for you here. give it the spec and then
- 19:10give it this prompt and then just start
- 19:11asking questions about AI coding like
- 19:13how I use sub aents uh what is the piv
- 19:16loop like just start asking and and
- 19:18seeing how easy it is for your agent to
- 19:20grab those things for you and so that's
- 19:22everything that I got for you today on
- 19:24OKF really is the future of personal
- 19:26agents if you appreciated this video
- 19:29you're looking forward to more things on
- 19:30AI coding and second brains I'd really
- 19:32appreciate a like and a subscribe and
- 19:34with that I will see you in the next
- 19:36video.
About this transcript
This page contains the full transcript of Finally, an Open Standard for the Karpathy LLM Wiki is HERE by Cole Medin, generated from the public captions YouTube serves with the video. The transcript has 4,105 words across 569 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.