God has shown the path to Auto Research!
If you're starting work in this area, check out the latest AI Scientist paper (TMLR2026)!
Our paper explores Auto Research from a similar perspective and helps clarify promising research directions and future work!
Jr. AI Scientist
Quote
Andrej Karpathy
@karpathy
·
I packaged up the "autoresearch" project into a new self-contained minimal repo if people would like to play over the weekend. It's basically nanochat LLM training core stripped down to a single-GPU, one file version of ~630 lines of code, then:
- the human iterates on the
Imagine web agents that don’t just browse but handle your tedious digital chores!
Our team developed WebChoreArena
- 532 human-curated tasks, crafted over 300+ hours
- Tests agents on massive information memorization, mathematical reasoning, and long-term memory
-
Ready for the next stage of multi-lingual LMM?
Happy to share our JMMMU, a Japanese MMMU benchmark!
For many users, it’s important to accelerate non-English research.
JMMMU will accelerate research in Japanese and multi-lingual LMMs!
HP: https://mmmu-japanese-benchmark.github.io/JMMMU/
Our vision for AGI is unlike the mainstream for the community.
Yes — we aim to build a super-human AI manga assistant
As the first step, our team developed MangaLMM, a LMM that can solve both MangaOCR and our newly created MangaVQA tasks!
https://arxiv.org/abs/2505.20298
Multiple-choice questions are a common benchmark format, but do LMMs really understand the answers?
In our #ACL2025 (main) paper, we propose Unsolvable Problem Detection to assess the robustness of LMM understanding!
Check it outhttps://arxiv.org/abs/2403.20331
Our survey on how OOD detection & related tasks have evolved in the VLM and Large VLM era is accepted to #TMLR!
The field is finally coming together, and OOD detection & anomaly detection are now at the center in the VLM era.
In the LVLM era, UPD (Unsolvable Problem
Ready for the next stage of multi-lingual LMM?
Happy to share our JMMMU, a Japanese MMMU benchmark!
For many users, it’s important to accelerate non-English research.
JMMMU will accelerate research in Japanese and multi-lingual LMMs!
HP: https://mmmu-japanese-benchmark.github.io/JMMMU/
How are OOD Detection, Open-set Recognition, Anomaly Detection, etc. evolving in the CLIP & GPT-4V eras? Check our preprint survey on Generalized OOD detection v2!
https://arxiv.org/abs/2407.21794
We encapsulate the evolution of these tasks in VLM era, identifying challenges!
April Fool is the birthday of Unsolvable Problem Detection! UPD examines the VLM’s ability to withhold answers when faced with unsolvable problems. Please enjoy VLMs with unsolvable problems today!
paper page: https://arxiv.org/abs/2403.20331
code: https://github.com/AtsuMiyai/UPD
Quote
AK
@_akhaliq
·
Unsolvable Problem Detection
Evaluating Trustworthiness of Vision Language Models
This paper introduces a novel and significant challenge for Vision Language Models (VLMs), termed Unsolvable Problem Detection (UPD). UPD examines the VLM's ability to withhold answers when
Our JMMMU has been accepted by #NAACL2025 main conference!
Kudos to the hard work of our great coauthors, reviewers, ACs, and many volunteers
Quote
Atsuyuki Miyai @UTokyo
@AtsuMiyaiAM
·
Ready for the next stage of multi-lingual LMM?
Happy to share our JMMMU, a Japanese MMMU benchmark!
For many users, it’s important to accelerate non-English research.
JMMMU will accelerate research in Japanese and multi-lingual LMMs!
HP: https://mmmu-japanese-benchmark.github.io/JMMMU/
Happy to share that our work "LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning" has been accepted to #NeurIPS2023!
I would like to express my deepest appreciation to my co-authors and reviewers.
arXiv:
(Just my personal opinion,)
I wish CV conferences could consider introducing a “Data-centric/Application Track,” similar to ICML or NeurIPS.
In my experience, CV conferences tend to equate novelty with proposing a specific approach to a specific problem, more so than in NLP or
I’ll attend #NeurIPS2024 and have an oral presentation@EvalEval Workshop 15th, 11:30 AM - 12:30 PM
Feel free to drop by if you can
Looking forward to exploring the intersection of diverse fields and catching up with the latest community trends!
Quote
Atsuyuki Miyai @UTokyo
@AtsuMiyaiAM
·
Ready for the next stage of multi-lingual LMM?
Happy to share our JMMMU, a Japanese MMMU benchmark!
For many users, it’s important to accelerate non-English research.
JMMMU will accelerate research in Japanese and multi-lingual LMMs!
HP: https://mmmu-japanese-benchmark.github.io/JMMMU/
will be there in person to present our JMMMU poster!
This is the very first step in one of my missions: bringing Japanese culture into work and sharing it with the world.
Come check it out!
Quote
Atsuyuki Miyai @UTokyo
@AtsuMiyaiAM
·
Ready for the next stage of multi-lingual LMM?
Happy to share our JMMMU, a Japanese MMMU benchmark!
For many users, it’s important to accelerate non-English research.
JMMMU will accelerate research in Japanese and multi-lingual LMMs!
HP: https://mmmu-japanese-benchmark.github.io/JMMMU/
MMUPD is now hosted on @huggingface Hub as a leaderboard
amazing to see LLaVA-1.6 is outperforming proprietary models on many subtasks
link in the next one! x.com/mervenoyann/st…
April Fool is the birthday of Unsolvable Problem Detection! UPD examines the VLM’s ability to withhold answers when faced with unsolvable problems. Please enjoy VLMs with unsolvable problems today!
paper page: https://arxiv.org/abs/2403.20331
code: https://github.com/AtsuMiyai/UPD
Thrilled that JMMMU has been featured on Yahoo News Japan!
https://news.yahoo.co.jp/articles/3887dcc709f016113968f628ca89f98057555772…
In the long run for multilingual LMMs, we seem to have established a strong starting point on the Japanese side.
This goal requires collaborative efforts. Let’s contribute together!
We update our #NeurIPS2023, “LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning”. LoCoOp performs OOD regularization with OOD regions (e.g., background) in ID training images and outperforms existing methods even in a 1-shot setting on ImageNet benchmarks.