YouTube2Text

Jia Li - A demo of Numina Studio — Transcript

by Institut des Hautes Etudes Scientifiques (IHES) · 3,633 words · 488 segments · language en · Watch on YouTube

Full transcript

  1. 0:12Hello everyone. Uh it's nice to be here
  2. 0:15and thank you for the invitation. Um so
  3. 0:18yeah, before the the screen shows up, I
  4. 0:20will say a few words about Luminina.
  5. 0:22Luminina is a nonprofit uh open source
  6. 0:24organization. We do we try to uh do our
  7. 0:29best to promote the use of AI in the
  8. 0:31field of mathematic and more general in
  9. 0:32fundamental science. Um and we're based
  10. 0:36in Paris. We have a small team here and
  11. 0:38we have done multiple works in the past
  12. 0:41really focusing on how to uh improve AI
  13. 0:44models for for for them to be able to do
  14. 0:48better mathematics. I will show you some
  15. 0:50result but today I'm mainly showing a
  16. 0:53preview of our work called Luminina
  17. 0:56Studio which is a SAS platform where you
  18. 0:59can use some or I'm I'm not sure how I'm
  19. 1:02supposed to show the screen
  20. 1:05is it am I doing something wrong? No.
  21. 1:07Anyway, uh so it's a platform where you
  22. 1:10will be able to orchestrate or or use
  23. 1:12orchestrated agent to solve uh uh
  24. 1:16complicated scientific problem
  25. 1:18especially uh informal theorem proving
  26. 1:20uh uh during very long time and we have
  27. 1:23been tested it with some mathematician
  28. 1:25to try to solve some non-trivial uh open
  29. 1:28problems. So uh very few words about
  30. 1:31Luminina. It's nonprofit. It's uh open
  31. 1:34source. Everything we have done or we uh
  32. 1:36we will do will be open source to the
  33. 1:38whole community. Uh it was founded uh in
  34. 1:41in in Paris uh by a couple of guys who
  35. 1:44are passionate about LM and mathematics.
  36. 1:47Um in the past we have been focusing on
  37. 1:50AI for or LM for mathematics. So we have
  38. 1:53been building some data sets, some
  39. 1:55models to help the community or the or
  40. 1:58or or the area to grow. Um a quick
  41. 2:02overview of what we have been we have
  42. 2:04done in the past. Uh we in the past when
  43. 2:07when like two or three years ago the
  44. 2:09focus is still uh helping this model to
  45. 2:12solve uh a competition level problem. So
  46. 2:15we started with a competition called AI
  47. 2:18uh MO which is uh the goal was to uh
  48. 2:22train the best AI to solve IMO level I
  49. 2:26mean pre IMO level problem let's say and
  50. 2:28we won the first prize we have released
  51. 2:30some data set to help uh and and back
  52. 2:32the time people are not aware or or
  53. 2:35people are still exploring how to how to
  54. 2:37train this model uh and then we received
  55. 2:40some funding and then in 20 last year we
  56. 2:42focused on formal mathematics uh uh the
  57. 2:45area was growing very fast and and and
  58. 2:47and and mainly driven by deep mine uh uh
  59. 2:50uh um uh with their team here and um and
  60. 2:54at the end of the 2025 we switch a bit
  61. 2:57our focus to build tools to help
  62. 2:59scientists uh to use AI to to solve
  63. 3:02challenging problem. So the platform
  64. 3:04that I'm going to show you very quickly
  65. 3:06and uh is a platform where you can
  66. 3:09orchestrate agent to help you to solve
  67. 3:12problem in
  68. 3:14very long to to to to really to to to to
  69. 3:17think how to make enable this uh uh LLM
  70. 3:21to be able to work as a human on very
  71. 3:23tough problem uh for a very long time.
  72. 3:26Um so one of in the same vein we have
  73. 3:29this work that we released like one or
  74. 3:30two month ago called luminina lean agent
  75. 3:32where we orchestrate uh uh LLM to to to
  76. 3:36do formal reasoning and uh well I mean
  77. 3:39back the time it's it was impressive to
  78. 3:41solve all the punam problem uh using
  79. 3:44math li and
  80. 3:46of course the the the area have been
  81. 3:49moving very fast uh uh a quick uh
  82. 3:52parenthesis uh we are also launching
  83. 3:54something what we called Luminina
  84. 3:56Fellowship where we work very closely
  85. 3:58with scientists and lab uh to design
  86. 4:02dedicated AI tools to help them to solve
  87. 4:04some problem that they believe AI could
  88. 4:06play a important role. Uh the first wave
  89. 4:10is finished but if you have idea or a
  90. 4:14great project that you think you believe
  91. 4:16AI could play play a critical role uh
  92. 4:19please reach out and we can discuss how
  93. 4:21we can might be able to collaborate
  94. 4:22together. So that's all a bit for the uh
  95. 4:25uh uh uh the presentation for I mean the
  96. 4:27story about Luminina. We're going to
  97. 4:29walk you through about our tool that we
  98. 4:31have been building since last two or
  99. 4:32three months. Uh the goal is to have a
  100. 4:35platform where you can uh enhance the
  101. 4:38LLM that you are using probably right
  102. 4:40now to build a team of agents so that
  103. 4:42you can have them work very long time on
  104. 4:45a very complicated uh problem starting
  105. 4:47with informal math reasoning or theory
  106. 4:50and proving. Uh so the idea is to say
  107. 4:53when you use uh chat GPT or even uh uh
  108. 4:57the pro version it's going to reason for
  109. 4:59a long time but it's going to be 10
  110. 5:01minute 30 minutes 1 hour here what we
  111. 5:04want to do is to say uh I mean a step
  112. 5:07forward toward maybe what we called
  113. 5:09autonomous research in the future is to
  114. 5:12say okay uh how can we orchestrate a
  115. 5:14bunch of agent to have them work on very
  116. 5:17tough problem and have my some change uh
  117. 5:20the and might have a chance to solve it
  118. 5:23at least give some good insight and have
  119. 5:25them work for a very long time. uh so uh
  120. 5:29using well based on this idea we build
  121. 5:31this platform that I'm going to do a
  122. 5:32very quick demo uh right after but on
  123. 5:35which you are you will be able to create
  124. 5:37a project uh describe what you want to
  125. 5:40do and then uh there is some well we
  126. 5:43have developed some by default hardness
  127. 5:46uh meaning that uh uh environment for
  128. 5:48the agent to work so that they can
  129. 5:50discuss brainstorm and uh uh uh uh come
  130. 5:53up with different approach and try to
  131. 5:55iterate on the proof until uh they the
  132. 5:59model believe uh uh uh finding something
  133. 6:01meaningful.
  134. 6:02Um
  135. 6:04and yeah today it's well it's it's we
  136. 6:08try to still focus on some uh uh use
  137. 6:11case like uh we have let's say three
  138. 6:14dedicated use case. The first one is
  139. 6:16scientific report to do literature
  140. 6:17search. The second is import in informal
  141. 6:20theorem proving and the third one is
  142. 6:22probably not relevant here. It's about
  143. 6:24uh simulation uh and for apply science.
  144. 6:28Um yeah, and the goal is really to have
  145. 6:31this agent work very long time on our
  146. 6:33platform. A task could last from minutes
  147. 6:35to days. Uh I have a back uh uh uh uh
  148. 6:38but the goal is that once you describe
  149. 6:40the task, you can leave and do something
  150. 6:42else and come back one day or two day
  151. 6:43later to see if the agent discover
  152. 6:45anything intelligent or not. So uh well
  153. 6:49I'm not an expert but uh even though you
  154. 6:51have my name on the paper but uh I'm
  155. 6:54have one uh here uh who has use our
  156. 6:57platform to solve one of the open
  157. 6:59problem that uh he encountered during
  158. 7:02his PhD uh and uh well I mean the the
  159. 7:07the and then he presented the final
  160. 7:08blueprint where he amended the the the
  161. 7:11the the proof of AI but
  162. 7:14as far as I know very few change has
  163. 7:16been added uh on the top of what a uh
  164. 7:19the model has has proposed. Uh so we
  165. 7:22have also a website today. We are in a
  166. 7:24private beta. So uh we will post a a
  167. 7:28message on Slack if you want to join and
  168. 7:31test. Probably in three or four weeks we
  169. 7:34will have a public release where people
  170. 7:35can really come and test. uh today. Uh
  171. 7:38yeah, and it's it probably it's going to
  172. 7:40be free for a relatively long period and
  173. 7:43the code will probably be open source uh
  174. 7:46a bit later.
  175. 7:47>> Yeah, I have a question.
  176. 7:49>> Yeah, go ahead. Sorry. Can you talk a
  177. 7:51bit about what is this notion of
  178. 7:53harness?
  179. 7:54>> Yes. Uh I mean let me I will do a quick
  180. 7:57demo uh later. So there is a live demo.
  181. 7:59I hope it's going to work. uh but and uh
  182. 8:03yeah the hard it's it's oh I mean there
  183. 8:06there have been been multiple period
  184. 8:08where where where people try to make LLM
  185. 8:12better. So in the very beginning we've
  186. 8:14realized or at least people realize if
  187. 8:16you do pump engineering that's uh
  188. 8:19meaning that uh the instruction that you
  189. 8:21give to the model could make a huge
  190. 8:23difference. So in the beginning we do
  191. 8:25pump engineering and then we do and we
  192. 8:28are switching to what we call today a
  193. 8:30lot of people are doing is called
  194. 8:32harness engineering. So the goal is it's
  195. 8:35always about context management. So it's
  196. 8:38really about uh uh like for example I
  197. 8:42mean I will take a quick example because
  198. 8:44uh it seems a bit abstract. So like in
  199. 8:47the harness that we proposed I mean we
  200. 8:48didn't invent this uh again the harness
  201. 8:52that we proposed here is very similar to
  202. 8:55a paper uh in 2025 called IMO Gemini
  203. 8:59agent something like that they built on
  204. 9:01top of back the time the Gemini version
  205. 9:05uh of developed by the mine to solve IMO
  206. 9:09level problem and they realized that
  207. 9:11they managed to reach a gold medal level
  208. 9:13using the the the back the time current
  209. 9:16version of Gemini. So what they have
  210. 9:18done is like instead of give a problem
  211. 9:20to Gemini and ask for the response they
  212. 9:24say okay and we use the same principle
  213. 9:26here. So we are going to have um a
  214. 9:29generator it's who is going to generate
  215. 9:32the proof and then you also have a
  216. 9:35verifier which uh you empty the context
  217. 9:38so that it's not biased by the very
  218. 9:40generator. So it's going to verify line
  219. 9:43by line the proof. And it turns out that
  220. 9:45if you do this uh and and the verifier
  221. 9:48also output a score to say okay I find
  222. 9:50some gaps in the reasoning there is a
  223. 9:52there is a grading scheme schema uh that
  224. 9:54allow the verifier to to to identify uh
  225. 9:58uh uh gaps in the in the in in the
  226. 10:00proof. It it turns out that if you do
  227. 10:02this iteratively uh multiple time uh the
  228. 10:06quality of the proof is much better than
  229. 10:08just calling uh Gemini one time. But
  230. 10:10this of course is very timeconuming. I
  231. 10:12have launched this task uh since this
  232. 10:14morning one hour. You can see what each
  233. 10:17agent has been doing. I'm not able to
  234. 10:19interpret it of course and you have the
  235. 10:20overall plan. Uh so the and and I don't
  236. 10:23even have the first version of the first
  237. 10:25proof. So like for example if you I mean
  238. 10:29some old uh if I take some I don't know
  239. 10:32if this version yeah like here the
  240. 10:35display is not great because we have
  241. 10:37fixed some bugs since then. But you can
  242. 10:38say that we you can see that we propose
  243. 10:40multiple version of the proof and each
  244. 10:43time the agent try to or the generator
  245. 10:45try to be a bit better. So the the to
  246. 10:48summarize the the the harness a bit like
  247. 10:51uh uh the the process or the workflow
  248. 10:54that you want the agent to follow so
  249. 10:57that it might produce something
  250. 11:00meaningfully better than just open up a
  251. 11:02chat GBT and just ask the question.
  252. 11:06That's basically
  253. 11:07>> and this workflow is something that you
  254. 11:09give beforehand or is it something you
  255. 11:11ask the model to figure it out?
  256. 11:13>> Uh today we have some default hardness.
  257. 11:16So if you open up a new project uh you
  258. 11:18say okay I'm going to do informal
  259. 11:20reasoning by default we use the IMO
  260. 11:23Gemini IMO agent uh uh to help you to do
  261. 11:26this. But we we will bring also some
  262. 11:28improvement. Uh so one core concept here
  263. 11:33which is quite popular recently is that
  264. 11:34you can if you have a data set you will
  265. 11:37be able to evolve the harness so that
  266. 11:40it's uh in so now what we are
  267. 11:43implementing is also evolving harness so
  268. 11:46the harness also become better so it
  269. 11:48means that the way that you orchestrate
  270. 11:50this LLM uh could become better and
  271. 11:53better uh when you collect a large
  272. 11:56enough data set of tough tough problem.
  273. 11:58So now you can see the agent working and
  274. 12:02then you have different generator
  275. 12:03verifier but I think the still the first
  276. 12:05loop is not finished yet. So it's going
  277. 12:07to I mean it the purpose of this
  278. 12:10platform is not for you to look at what
  279. 12:12the model do all the time uh but rather
  280. 12:16than you have some task you p uh you you
  281. 12:19you you send them to the agent and then
  282. 12:22you go to something else and then you
  283. 12:24come back one day two day to collect
  284. 12:26what you get. So that's kind of the
  285. 12:27admission platform.
  286. 12:29>> Is the verifier trained in a different
  287. 12:31way than the other agents? Is it trained
  288. 12:32on Lynn or doing some formal?
  289. 12:34>> Yes, we have uh formal platform today.
  290. 12:38No, the short answer is not.
  291. 12:40>> Yeah, the verifier is trained the same
  292. 12:42fashion. I mean I didn't train the
  293. 12:44model. Uh we we are using GPT 5.5 as
  294. 12:47verifier today. Uh and and and we have
  295. 12:50two generator. One is GPT 5.5 and the
  296. 12:52other one is Gemini uh 3.5 flesh. I
  297. 12:55think I I don't know if it is 3.5 or
  298. 12:58still 3.1 but anyway uh today I think
  299. 13:00the verifier I mean inside I don't have
  300. 13:04all the information from open AAI but I
  301. 13:06think they train gener generator and
  302. 13:08verifier within the same model. Yeah,
  303. 13:10but of course we can if we I mean I I do
  304. 13:15think the the formal reasoning still
  305. 13:17have some gaps to catch up. But in the
  306. 13:20future maybe the verifier could be
  307. 13:22formal but today I don't I don't
  308. 13:26with the compute that we have I don't
  309. 13:28think a formal verifier will bring
  310. 13:30meaningful insight to the model.
  311. 13:33>> But when you say it's a verifier it also
  312. 13:35give a score to the proof. Yes.
  313. 13:37>> Yes. Yes. Yes. Yes. So it's a verifier
  314. 13:40score so that you can and you prompt it
  315. 13:43to have a grading schema to say okay if
  316. 13:45there's a reasoning gap you're going to
  317. 13:48minus 10 and then if there is other bugs
  318. 13:51if there is something that uh it's
  319. 13:53completely false and then you minus 50
  320. 13:55something like that. So if the verifier
  321. 13:57and the generator train in the same way
  322. 13:59why the generator cannot identify the
  323. 14:0140s themselves.
  324. 14:03>> Yes, that's a mystery. I mean there was
  325. 14:05some explanation like for example why I
  326. 14:08mean today it's less and less the case
  327. 14:10but in the very beginning when you have
  328. 14:12a generator and then you you just ask
  329. 14:16the generator can you verify the proof
  330. 14:19uh the generator get bias because
  331. 14:23in the context you get all the reasoning
  332. 14:26step to get to the proof. So in the very
  333. 14:28beginning people realize that if you
  334. 14:31clean the context you you remove the
  335. 14:33reasoning traces of the generator and
  336. 14:35ask a new agent to say okay please
  337. 14:37verify this proof it's performing much
  338. 14:39better than asking the generator itself
  339. 14:41to verify the proof. I think today LLM
  340. 14:44are more capable. So this case uh this
  341. 14:46become less true. But uh the the
  342. 14:49intuition behind is like uh train the
  343. 14:52model to verify something. It's much
  344. 14:55easier to come up with a complete proof.
  345. 14:57So that's kind of the inside bit high
  346. 15:00but you I I do imagine I mean today
  347. 15:02there will be less and less literature
  348. 15:04about how to train this kind of
  349. 15:05generator verifier because all this
  350. 15:07company they are clos but the closest
  351. 15:10paper that I mean give you a lot of uh I
  352. 15:14mean training inside about how this
  353. 15:15model are trained is the DC V2 paper and
  354. 15:18they do have like adversary training
  355. 15:21loop. So they they they make the
  356. 15:22verifier stronger and then they make the
  357. 15:25generator stronger. So they they try to
  358. 15:27do this to make both better but but it
  359. 15:30seems like the the why this work is
  360. 15:32because verifier are easier to train.
  361. 15:36Can
  362. 15:37>> the system keep a log of the
  363. 15:39interactions between the agents for
  364. 15:41later?
  365. 15:42>> Yes. So normally uh you will be you are
  366. 15:46you will be able to download everything.
  367. 15:48uh we are we are finding a way to
  368. 15:51visualize them like uh a Wikipedia of
  369. 15:54how the proof have been generated and
  370. 15:56all the exploration. So we might have an
  371. 15:58upgrade in one or two days but today
  372. 16:01what you will be able to do is to uh I
  373. 16:05think yeah well I mean I have a this
  374. 16:07really dumb example of uh infinite many
  375. 16:11prime. So if you go to uh I think I
  376. 16:14launched this this morning. So you have
  377. 16:15the conclusion, the summary. Well, I
  378. 16:17mean that that you you got one proof. Uh
  379. 16:19I think it two proof is generated. So
  380. 16:22you can download all the reasoning
  381. 16:24traces and then I can have like this is
  382. 16:27the first version. Uh I don't know if I
  383. 16:30have the criticism.
  384. 16:32No, I don't have the criticism of the so
  385. 16:35yeah this is the complete uh version of
  386. 16:37the proof. I I don't have we we will
  387. 16:39record the in in the future version we
  388. 16:41will record the criticism of the
  389. 16:42verifier as well so that you can see why
  390. 16:45the model believe it's not a good proof
  391. 16:48but I think this is too easy so that you
  392. 16:50didn't need criticism to to make a
  393. 16:52perfect proof
  394. 16:54but you do have the interaction we will
  395. 16:56try to make it more intuitive and easier
  396. 16:58to explore today what you get is a
  397. 17:02series of PDF where the generator
  398. 17:04proposed multiple version of the proof
  399. 17:06and then the verifier come back and
  400. 17:08create a size and uh and and and so that
  401. 17:10the generator could create a new
  402. 17:12version. So that's kind of a relatively
  403. 17:15simple workflow.
  404. 17:18Talking about specific uh features
  405. 17:22something which is very useful in this
  406. 17:23business is to be able to produce a
  407. 17:25quick related
  408. 17:27paper and so
  409. 17:29>> ah yes
  410. 17:31>> yeah we can do that it's on our road
  411. 17:33map. Uh uh yeah.
  412. 17:35>> Uh looking at these traces, do you think
  413. 17:37the model knows what it's doing in the
  414. 17:39sense that it's proving a mathematical
  415. 17:41statement or something or is it just
  416. 17:43like just so that it has become so good
  417. 17:46at memorizing patterns or whatever it's
  418. 17:48just outputting? Can you distinguish
  419. 17:50between these like is is the me like
  420. 17:53what I'm asking is the model like a very
  421. 17:54good I don't know mathematical memorizer
  422. 17:57so that it can spit out these proofs
  423. 17:59just like it's seconds or a model that
  424. 18:02understands what it's doing in the sense
  425. 18:04like it made a mistake it goes back I
  426. 18:06know this is a bit hard to what do you
  427. 18:08say
  428. 18:09>> yeah I think in general now model are
  429. 18:11very capable so we don't study this but
  430. 18:13I think in general model are able to
  431. 18:15generalize before beyond just memorizing
  432. 18:17I think I'm a professor am yet he has
  433. 18:21multiple study like uh studying the
  434. 18:23lamas that have been published last year
  435. 18:25on archive last week sorry and this
  436. 18:27shows that this model are really able to
  437. 18:28generalize and progress so I I I do
  438. 18:31think I mean sometimes they do memorize
  439. 18:32stuff but in general they will be able
  440. 18:34to go beyond their training set yeah we
  441. 18:38don't train model anymore well for a
  442. 18:40while so but maybe in the future we will
  443. 18:42we will but but but I think today for
  444. 18:45people who do train model they do
  445. 18:46realize that this people this model can
  446. 18:48generalize much better than than before.
  447. 18:51>> Yeah, please.
  448. 18:53>> So, you have the smaller model that can
  449. 18:55run on the computer, right?
  450. 18:57>> Yeah.
  451. 18:58>> So, what can you do with this type of
  452. 19:01models?
  453. 19:02>> Um, unfortunately, if you don't want to
  454. 19:05do frontier science, it's very
  455. 19:07unfortunately you I don't see a way that
  456. 19:10you can do it with the models on your
  457. 19:11computer. I mean you can do a lot of
  458. 19:13stuff like helping you to uh do
  459. 19:16literature search automatic automatize
  460. 19:19workflow but to do frontier science
  461. 19:22unfortunately model uh we see a
  462. 19:24important and more and more important
  463. 19:26gap between uh top tier company like
  464. 19:29Google demand and open AI versus uh
  465. 19:32their open source uh counterpart uh and
  466. 19:34not mentioning about small small models.
  467. 19:36Yeah.
  468. 19:38So I don't think
  469. 19:39>> but I mean like if you give a proof a
  470. 19:42short proof of a little technical
  471. 19:46simple scientific I would say can it
  472. 19:49check the proof
  473. 19:51>> for small model I doubt that uh it's
  474. 19:53really it's really model size really
  475. 19:55count when you have difficult problem
  476. 19:56for easy problem like high high school
  477. 19:59level I think it's capable but for
  478. 20:00research level problem I think you will
  479. 20:02have to I mean one way or another use
  480. 20:04this frontier model from demine and open
  481. 20:06AAI
  482. 20:12Okay. Well, I mean if you're interested
  483. 20:14uh to try and if you have a project
  484. 20:16where you think AI can make a
  485. 20:18difference, feel free to reach out and I
  486. 20:20think we have posted something on Zulib
  487. 20:22and so feel free to interact as well and
  488. 20:25thank you.

About this transcript

This page contains the full transcript of Jia Li - A demo of Numina Studio by Institut des Hautes Etudes Scientifiques (IHES), generated from the public captions YouTube serves with the video. The transcript has 3,633 words across 488 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.