1. X
  2. Haruki Nishimura
Log inSign up
Haruki Nishimura
504 posts
Image
user avatar
Haruki Nishimura
@imp_aa
Learning and planning for safe, embodied autonomous systems under uncertainty. Senior Research Scientist @ToyotaResearch. PhD from @StanfordMSL. 日本語 & English
California, USA
harukins.github.io
Joined March 2018
763
Following
700
Followers
RepliesRepliesMediaMedia
  • user avatar
    Haruki Nishimura
    @imp_aa
    Jun 19
    Sample-efficient and reliable policy comparison is essential for both fast design iteration and reproducible baseline comparison, where expensive hardware eval remains the gold standard. Check out this awesome RSS work led by @das_princeton on a new, general framework!
    user avatar
    David Snyder
    @das_princeton
    Jun 15
    (1/12) How should we rigorously compare robot policies? Comparison is central to robotics research, but is inherently expensive. We introduce NSCORE, a flexible, general-purpose, and data-efficient method for rigorous policy comparison. Accepted to RSS 2026.
    Image
    1.1K
  • user avatar
    Haruki Nishimura
    @imp_aa
    Jun 19
    Robot evaluation is an open problem, especially in the age of foundation models. Check out this blog post sharing actionable insights from our RSS Workshop last year, highlighting challenges and describing best practices. Huge thanks to @hocherie1 for leading this effort!
    user avatar
    Cherie Ho
    @hocherie1
    Jun 18
    How should we evaluate robots in the age of foundation models? We hosted the RSS RoboEval Workshop with folks from academia, industry, and policy to discuss this. 💪 We put together an actionable guide and insights to get you started on robot evaluations. Link below. 🧵1/N
    Image
    436
  • user avatar
    Haruki Nishimura
    @imp_aa
    Apr 22
    This is hugely based on @das_princeton's implementation that came out of the collaboration between TLU tri.global/trustworthy-le… and @Majumdar_Ani's group at Princeton out of an internship project!
    user avatar
    Katherine Liu
    @robo_kat
    Apr 22
    This is actually a pretty big deal — we rely on @imp_aa’s implementations to tell when policies are statistically different than each other. If someone presents some quick mean-only results internally without the CLD analysis, you can be sure someone will eventually ask for it.
    821
  • user avatar
    Haruki Nishimura
    @imp_aa
    Apr 22
    A huge shout-out to TRI's VLA team for the public release of VLA Foundry! You can take full control of VLA training with this fully open-sourced codebase, which comes with a nice GUI dashboard with rigorous policy comparison powered by STEP🪜 tri-ml.github.io/step/
    user avatar
    Jean Mercat
    @MercatJean
    Apr 22
    Releasing VLA Foundry: an open-source framework that unifies LLM, VLM, and VLA training in a single codebase. End-to-end control from language pretraining to action-expert fine-tuning — no more stitching together incompatible repos.
    Image
    00:00
    7.8K
  • user avatar
    Haruki Nishimura
    @imp_aa
    Apr 13
    Congrats to the @LeRobotHF team on this remarkable contribution to the robotics community by open-sourcing "everything" including code, data, and all the valuable knowledge! Our TLU team at TRI is fortunate to have collaborated on statistical evaluation and analysis.
    user avatar
    LeRobot
    @LeRobotHF
    Apr 7
    Releasing the Unfolding Robotics blog! Time to unfold robotics: we trained a robot to fold clothes using 8 bimanual setups, 100+ hours of demonstrations, and 5k+ GPU hours. Flashy robot demos are everywhere. But you rarely see the real story: the data, the failures, the
    Image
    00:00
    916
  • See @imp_aa's full profile

    Sign up
    Log in

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement