Jia Li - A demo of Numina Studio — Transcript
Full transcript
- 0:12Hello everyone. Uh it's nice to be here
- 0:15and thank you for the invitation. Um so
- 0:18yeah, before the the screen shows up, I
- 0:20will say a few words about Luminina.
- 0:22Luminina is a nonprofit uh open source
- 0:24organization. We do we try to uh do our
- 0:29best to promote the use of AI in the
- 0:31field of mathematic and more general in
- 0:32fundamental science. Um and we're based
- 0:36in Paris. We have a small team here and
- 0:38we have done multiple works in the past
- 0:41really focusing on how to uh improve AI
- 0:44models for for for them to be able to do
- 0:48better mathematics. I will show you some
- 0:50result but today I'm mainly showing a
- 0:53preview of our work called Luminina
- 0:56Studio which is a SAS platform where you
- 0:59can use some or I'm I'm not sure how I'm
- 1:02supposed to show the screen
- 1:05is it am I doing something wrong? No.
- 1:07Anyway, uh so it's a platform where you
- 1:10will be able to orchestrate or or use
- 1:12orchestrated agent to solve uh uh
- 1:16complicated scientific problem
- 1:18especially uh informal theorem proving
- 1:20uh uh during very long time and we have
- 1:23been tested it with some mathematician
- 1:25to try to solve some non-trivial uh open
- 1:28problems. So uh very few words about
- 1:31Luminina. It's nonprofit. It's uh open
- 1:34source. Everything we have done or we uh
- 1:36we will do will be open source to the
- 1:38whole community. Uh it was founded uh in
- 1:41in in Paris uh by a couple of guys who
- 1:44are passionate about LM and mathematics.
- 1:47Um in the past we have been focusing on
- 1:50AI for or LM for mathematics. So we have
- 1:53been building some data sets, some
- 1:55models to help the community or the or
- 1:58or or the area to grow. Um a quick
- 2:02overview of what we have been we have
- 2:04done in the past. Uh we in the past when
- 2:07when like two or three years ago the
- 2:09focus is still uh helping this model to
- 2:12solve uh a competition level problem. So
- 2:15we started with a competition called AI
- 2:18uh MO which is uh the goal was to uh
- 2:22train the best AI to solve IMO level I
- 2:26mean pre IMO level problem let's say and
- 2:28we won the first prize we have released
- 2:30some data set to help uh and and back
- 2:32the time people are not aware or or
- 2:35people are still exploring how to how to
- 2:37train this model uh and then we received
- 2:40some funding and then in 20 last year we
- 2:42focused on formal mathematics uh uh the
- 2:45area was growing very fast and and and
- 2:47and and mainly driven by deep mine uh uh
- 2:50uh um uh with their team here and um and
- 2:54at the end of the 2025 we switch a bit
- 2:57our focus to build tools to help
- 2:59scientists uh to use AI to to solve
- 3:02challenging problem. So the platform
- 3:04that I'm going to show you very quickly
- 3:06and uh is a platform where you can
- 3:09orchestrate agent to help you to solve
- 3:12problem in
- 3:14very long to to to to really to to to to
- 3:17think how to make enable this uh uh LLM
- 3:21to be able to work as a human on very
- 3:23tough problem uh for a very long time.
- 3:26Um so one of in the same vein we have
- 3:29this work that we released like one or
- 3:30two month ago called luminina lean agent
- 3:32where we orchestrate uh uh LLM to to to
- 3:36do formal reasoning and uh well I mean
- 3:39back the time it's it was impressive to
- 3:41solve all the punam problem uh using
- 3:44math li and
- 3:46of course the the the area have been
- 3:49moving very fast uh uh a quick uh
- 3:52parenthesis uh we are also launching
- 3:54something what we called Luminina
- 3:56Fellowship where we work very closely
- 3:58with scientists and lab uh to design
- 4:02dedicated AI tools to help them to solve
- 4:04some problem that they believe AI could
- 4:06play a important role. Uh the first wave
- 4:10is finished but if you have idea or a
- 4:14great project that you think you believe
- 4:16AI could play play a critical role uh
- 4:19please reach out and we can discuss how
- 4:21we can might be able to collaborate
- 4:22together. So that's all a bit for the uh
- 4:25uh uh uh the presentation for I mean the
- 4:27story about Luminina. We're going to
- 4:29walk you through about our tool that we
- 4:31have been building since last two or
- 4:32three months. Uh the goal is to have a
- 4:35platform where you can uh enhance the
- 4:38LLM that you are using probably right
- 4:40now to build a team of agents so that
- 4:42you can have them work very long time on
- 4:45a very complicated uh problem starting
- 4:47with informal math reasoning or theory
- 4:50and proving. Uh so the idea is to say
- 4:53when you use uh chat GPT or even uh uh
- 4:57the pro version it's going to reason for
- 4:59a long time but it's going to be 10
- 5:01minute 30 minutes 1 hour here what we
- 5:04want to do is to say uh I mean a step
- 5:07forward toward maybe what we called
- 5:09autonomous research in the future is to
- 5:12say okay uh how can we orchestrate a
- 5:14bunch of agent to have them work on very
- 5:17tough problem and have my some change uh
- 5:20the and might have a chance to solve it
- 5:23at least give some good insight and have
- 5:25them work for a very long time. uh so uh
- 5:29using well based on this idea we build
- 5:31this platform that I'm going to do a
- 5:32very quick demo uh right after but on
- 5:35which you are you will be able to create
- 5:37a project uh describe what you want to
- 5:40do and then uh there is some well we
- 5:43have developed some by default hardness
- 5:46uh meaning that uh uh environment for
- 5:48the agent to work so that they can
- 5:50discuss brainstorm and uh uh uh uh come
- 5:53up with different approach and try to
- 5:55iterate on the proof until uh they the
- 5:59model believe uh uh uh finding something
- 6:01meaningful.
- 6:02Um
- 6:04and yeah today it's well it's it's we
- 6:08try to still focus on some uh uh use
- 6:11case like uh we have let's say three
- 6:14dedicated use case. The first one is
- 6:16scientific report to do literature
- 6:17search. The second is import in informal
- 6:20theorem proving and the third one is
- 6:22probably not relevant here. It's about
- 6:24uh simulation uh and for apply science.
- 6:28Um yeah, and the goal is really to have
- 6:31this agent work very long time on our
- 6:33platform. A task could last from minutes
- 6:35to days. Uh I have a back uh uh uh uh
- 6:38but the goal is that once you describe
- 6:40the task, you can leave and do something
- 6:42else and come back one day or two day
- 6:43later to see if the agent discover
- 6:45anything intelligent or not. So uh well
- 6:49I'm not an expert but uh even though you
- 6:51have my name on the paper but uh I'm
- 6:54have one uh here uh who has use our
- 6:57platform to solve one of the open
- 6:59problem that uh he encountered during
- 7:02his PhD uh and uh well I mean the the
- 7:07the and then he presented the final
- 7:08blueprint where he amended the the the
- 7:11the the proof of AI but
- 7:14as far as I know very few change has
- 7:16been added uh on the top of what a uh
- 7:19the model has has proposed. Uh so we
- 7:22have also a website today. We are in a
- 7:24private beta. So uh we will post a a
- 7:28message on Slack if you want to join and
- 7:31test. Probably in three or four weeks we
- 7:34will have a public release where people
- 7:35can really come and test. uh today. Uh
- 7:38yeah, and it's it probably it's going to
- 7:40be free for a relatively long period and
- 7:43the code will probably be open source uh
- 7:46a bit later.
- 7:47>> Yeah, I have a question.
- 7:49>> Yeah, go ahead. Sorry. Can you talk a
- 7:51bit about what is this notion of
- 7:53harness?
- 7:54>> Yes. Uh I mean let me I will do a quick
- 7:57demo uh later. So there is a live demo.
- 7:59I hope it's going to work. uh but and uh
- 8:03yeah the hard it's it's oh I mean there
- 8:06there have been been multiple period
- 8:08where where where people try to make LLM
- 8:12better. So in the very beginning we've
- 8:14realized or at least people realize if
- 8:16you do pump engineering that's uh
- 8:19meaning that uh the instruction that you
- 8:21give to the model could make a huge
- 8:23difference. So in the beginning we do
- 8:25pump engineering and then we do and we
- 8:28are switching to what we call today a
- 8:30lot of people are doing is called
- 8:32harness engineering. So the goal is it's
- 8:35always about context management. So it's
- 8:38really about uh uh like for example I
- 8:42mean I will take a quick example because
- 8:44uh it seems a bit abstract. So like in
- 8:47the harness that we proposed I mean we
- 8:48didn't invent this uh again the harness
- 8:52that we proposed here is very similar to
- 8:55a paper uh in 2025 called IMO Gemini
- 8:59agent something like that they built on
- 9:01top of back the time the Gemini version
- 9:05uh of developed by the mine to solve IMO
- 9:09level problem and they realized that
- 9:11they managed to reach a gold medal level
- 9:13using the the the back the time current
- 9:16version of Gemini. So what they have
- 9:18done is like instead of give a problem
- 9:20to Gemini and ask for the response they
- 9:24say okay and we use the same principle
- 9:26here. So we are going to have um a
- 9:29generator it's who is going to generate
- 9:32the proof and then you also have a
- 9:35verifier which uh you empty the context
- 9:38so that it's not biased by the very
- 9:40generator. So it's going to verify line
- 9:43by line the proof. And it turns out that
- 9:45if you do this uh and and the verifier
- 9:48also output a score to say okay I find
- 9:50some gaps in the reasoning there is a
- 9:52there is a grading scheme schema uh that
- 9:54allow the verifier to to to identify uh
- 9:58uh uh gaps in the in the in in the
- 10:00proof. It it turns out that if you do
- 10:02this iteratively uh multiple time uh the
- 10:06quality of the proof is much better than
- 10:08just calling uh Gemini one time. But
- 10:10this of course is very timeconuming. I
- 10:12have launched this task uh since this
- 10:14morning one hour. You can see what each
- 10:17agent has been doing. I'm not able to
- 10:19interpret it of course and you have the
- 10:20overall plan. Uh so the and and I don't
- 10:23even have the first version of the first
- 10:25proof. So like for example if you I mean
- 10:29some old uh if I take some I don't know
- 10:32if this version yeah like here the
- 10:35display is not great because we have
- 10:37fixed some bugs since then. But you can
- 10:38say that we you can see that we propose
- 10:40multiple version of the proof and each
- 10:43time the agent try to or the generator
- 10:45try to be a bit better. So the the to
- 10:48summarize the the the harness a bit like
- 10:51uh uh the the process or the workflow
- 10:54that you want the agent to follow so
- 10:57that it might produce something
- 11:00meaningfully better than just open up a
- 11:02chat GBT and just ask the question.
- 11:06That's basically
- 11:07>> and this workflow is something that you
- 11:09give beforehand or is it something you
- 11:11ask the model to figure it out?
- 11:13>> Uh today we have some default hardness.
- 11:16So if you open up a new project uh you
- 11:18say okay I'm going to do informal
- 11:20reasoning by default we use the IMO
- 11:23Gemini IMO agent uh uh to help you to do
- 11:26this. But we we will bring also some
- 11:28improvement. Uh so one core concept here
- 11:33which is quite popular recently is that
- 11:34you can if you have a data set you will
- 11:37be able to evolve the harness so that
- 11:40it's uh in so now what we are
- 11:43implementing is also evolving harness so
- 11:46the harness also become better so it
- 11:48means that the way that you orchestrate
- 11:50this LLM uh could become better and
- 11:53better uh when you collect a large
- 11:56enough data set of tough tough problem.
- 11:58So now you can see the agent working and
- 12:02then you have different generator
- 12:03verifier but I think the still the first
- 12:05loop is not finished yet. So it's going
- 12:07to I mean it the purpose of this
- 12:10platform is not for you to look at what
- 12:12the model do all the time uh but rather
- 12:16than you have some task you p uh you you
- 12:19you you send them to the agent and then
- 12:22you go to something else and then you
- 12:24come back one day two day to collect
- 12:26what you get. So that's kind of the
- 12:27admission platform.
- 12:29>> Is the verifier trained in a different
- 12:31way than the other agents? Is it trained
- 12:32on Lynn or doing some formal?
- 12:34>> Yes, we have uh formal platform today.
- 12:38No, the short answer is not.
- 12:40>> Yeah, the verifier is trained the same
- 12:42fashion. I mean I didn't train the
- 12:44model. Uh we we are using GPT 5.5 as
- 12:47verifier today. Uh and and and we have
- 12:50two generator. One is GPT 5.5 and the
- 12:52other one is Gemini uh 3.5 flesh. I
- 12:55think I I don't know if it is 3.5 or
- 12:58still 3.1 but anyway uh today I think
- 13:00the verifier I mean inside I don't have
- 13:04all the information from open AAI but I
- 13:06think they train gener generator and
- 13:08verifier within the same model. Yeah,
- 13:10but of course we can if we I mean I I do
- 13:15think the the formal reasoning still
- 13:17have some gaps to catch up. But in the
- 13:20future maybe the verifier could be
- 13:22formal but today I don't I don't
- 13:26with the compute that we have I don't
- 13:28think a formal verifier will bring
- 13:30meaningful insight to the model.
- 13:33>> But when you say it's a verifier it also
- 13:35give a score to the proof. Yes.
- 13:37>> Yes. Yes. Yes. Yes. So it's a verifier
- 13:40score so that you can and you prompt it
- 13:43to have a grading schema to say okay if
- 13:45there's a reasoning gap you're going to
- 13:48minus 10 and then if there is other bugs
- 13:51if there is something that uh it's
- 13:53completely false and then you minus 50
- 13:55something like that. So if the verifier
- 13:57and the generator train in the same way
- 13:59why the generator cannot identify the
- 14:0140s themselves.
- 14:03>> Yes, that's a mystery. I mean there was
- 14:05some explanation like for example why I
- 14:08mean today it's less and less the case
- 14:10but in the very beginning when you have
- 14:12a generator and then you you just ask
- 14:16the generator can you verify the proof
- 14:19uh the generator get bias because
- 14:23in the context you get all the reasoning
- 14:26step to get to the proof. So in the very
- 14:28beginning people realize that if you
- 14:31clean the context you you remove the
- 14:33reasoning traces of the generator and
- 14:35ask a new agent to say okay please
- 14:37verify this proof it's performing much
- 14:39better than asking the generator itself
- 14:41to verify the proof. I think today LLM
- 14:44are more capable. So this case uh this
- 14:46become less true. But uh the the
- 14:49intuition behind is like uh train the
- 14:52model to verify something. It's much
- 14:55easier to come up with a complete proof.
- 14:57So that's kind of the inside bit high
- 15:00but you I I do imagine I mean today
- 15:02there will be less and less literature
- 15:04about how to train this kind of
- 15:05generator verifier because all this
- 15:07company they are clos but the closest
- 15:10paper that I mean give you a lot of uh I
- 15:14mean training inside about how this
- 15:15model are trained is the DC V2 paper and
- 15:18they do have like adversary training
- 15:21loop. So they they they make the
- 15:22verifier stronger and then they make the
- 15:25generator stronger. So they they try to
- 15:27do this to make both better but but it
- 15:30seems like the the why this work is
- 15:32because verifier are easier to train.
- 15:36Can
- 15:37>> the system keep a log of the
- 15:39interactions between the agents for
- 15:41later?
- 15:42>> Yes. So normally uh you will be you are
- 15:46you will be able to download everything.
- 15:48uh we are we are finding a way to
- 15:51visualize them like uh a Wikipedia of
- 15:54how the proof have been generated and
- 15:56all the exploration. So we might have an
- 15:58upgrade in one or two days but today
- 16:01what you will be able to do is to uh I
- 16:05think yeah well I mean I have a this
- 16:07really dumb example of uh infinite many
- 16:11prime. So if you go to uh I think I
- 16:14launched this this morning. So you have
- 16:15the conclusion, the summary. Well, I
- 16:17mean that that you you got one proof. Uh
- 16:19I think it two proof is generated. So
- 16:22you can download all the reasoning
- 16:24traces and then I can have like this is
- 16:27the first version. Uh I don't know if I
- 16:30have the criticism.
- 16:32No, I don't have the criticism of the so
- 16:35yeah this is the complete uh version of
- 16:37the proof. I I don't have we we will
- 16:39record the in in the future version we
- 16:41will record the criticism of the
- 16:42verifier as well so that you can see why
- 16:45the model believe it's not a good proof
- 16:48but I think this is too easy so that you
- 16:50didn't need criticism to to make a
- 16:52perfect proof
- 16:54but you do have the interaction we will
- 16:56try to make it more intuitive and easier
- 16:58to explore today what you get is a
- 17:02series of PDF where the generator
- 17:04proposed multiple version of the proof
- 17:06and then the verifier come back and
- 17:08create a size and uh and and and so that
- 17:10the generator could create a new
- 17:12version. So that's kind of a relatively
- 17:15simple workflow.
- 17:18Talking about specific uh features
- 17:22something which is very useful in this
- 17:23business is to be able to produce a
- 17:25quick related
- 17:27paper and so
- 17:29>> ah yes
- 17:31>> yeah we can do that it's on our road
- 17:33map. Uh uh yeah.
- 17:35>> Uh looking at these traces, do you think
- 17:37the model knows what it's doing in the
- 17:39sense that it's proving a mathematical
- 17:41statement or something or is it just
- 17:43like just so that it has become so good
- 17:46at memorizing patterns or whatever it's
- 17:48just outputting? Can you distinguish
- 17:50between these like is is the me like
- 17:53what I'm asking is the model like a very
- 17:54good I don't know mathematical memorizer
- 17:57so that it can spit out these proofs
- 17:59just like it's seconds or a model that
- 18:02understands what it's doing in the sense
- 18:04like it made a mistake it goes back I
- 18:06know this is a bit hard to what do you
- 18:08say
- 18:09>> yeah I think in general now model are
- 18:11very capable so we don't study this but
- 18:13I think in general model are able to
- 18:15generalize before beyond just memorizing
- 18:17I think I'm a professor am yet he has
- 18:21multiple study like uh studying the
- 18:23lamas that have been published last year
- 18:25on archive last week sorry and this
- 18:27shows that this model are really able to
- 18:28generalize and progress so I I I do
- 18:31think I mean sometimes they do memorize
- 18:32stuff but in general they will be able
- 18:34to go beyond their training set yeah we
- 18:38don't train model anymore well for a
- 18:40while so but maybe in the future we will
- 18:42we will but but but I think today for
- 18:45people who do train model they do
- 18:46realize that this people this model can
- 18:48generalize much better than than before.
- 18:51>> Yeah, please.
- 18:53>> So, you have the smaller model that can
- 18:55run on the computer, right?
- 18:57>> Yeah.
- 18:58>> So, what can you do with this type of
- 19:01models?
- 19:02>> Um, unfortunately, if you don't want to
- 19:05do frontier science, it's very
- 19:07unfortunately you I don't see a way that
- 19:10you can do it with the models on your
- 19:11computer. I mean you can do a lot of
- 19:13stuff like helping you to uh do
- 19:16literature search automatic automatize
- 19:19workflow but to do frontier science
- 19:22unfortunately model uh we see a
- 19:24important and more and more important
- 19:26gap between uh top tier company like
- 19:29Google demand and open AI versus uh
- 19:32their open source uh counterpart uh and
- 19:34not mentioning about small small models.
- 19:36Yeah.
- 19:38So I don't think
- 19:39>> but I mean like if you give a proof a
- 19:42short proof of a little technical
- 19:46simple scientific I would say can it
- 19:49check the proof
- 19:51>> for small model I doubt that uh it's
- 19:53really it's really model size really
- 19:55count when you have difficult problem
- 19:56for easy problem like high high school
- 19:59level I think it's capable but for
- 20:00research level problem I think you will
- 20:02have to I mean one way or another use
- 20:04this frontier model from demine and open
- 20:06AAI
- 20:12Okay. Well, I mean if you're interested
- 20:14uh to try and if you have a project
- 20:16where you think AI can make a
- 20:18difference, feel free to reach out and I
- 20:20think we have posted something on Zulib
- 20:22and so feel free to interact as well and
- 20:25thank you.
About this transcript
This page contains the full transcript of Jia Li - A demo of Numina Studio by Institut des Hautes Etudes Scientifiques (IHES), generated from the public captions YouTube serves with the video. The transcript has 3,633 words across 488 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.