2027 Internship Evaluation Engineer, Metric Prototyping
San Francisco, CA - USA
Department:
Job Summary
At Bedrock were moving AI out of the lab and into the real world. Our team includes veterans who helped launch Waymo scaled Segment to a $3.2B acquisition and grew Uber Freight to $5B in revenue. Today were deploying autonomous systems on heavy construction equipment across the country improving safety on job sites and accelerating schedules on critical infrastructure projects.
Were not here debating the future of AI. Were deploying it in the real just two years weve raised $350M and achieved the first fully autonomous excavator deployments in construction.
This is where algorithms meet steel-toed boots. Youll work alongside construction veterans and world-class engineers to solve physical-world problems that simulations cant touch. If youre ready to do meaningful work on hard problems wed love to have you join us.
Every time Bedrock works with a new machine for tasks like grading pads digging trenches or loading trucks we need a way to verify the machine is actually performing well. The Evaluation team defines what good looks like: building the metrics benchmarks and evaluation pipelines that tell us whether our autonomy stack is improving or just changing. As our Evaluation intern youll prototype new metrics for future machine tasks the ones we havent shipped yet so that when the autonomy team is ready to train theres already a clear scoreboard.
Partner closely with data understanding and ML engineers to rapidly iterate on new metric formulations and task benchmarks
Mine fleet datasets for key reference examples and edge cases then test proposed metrics against live runs and historical logs
Get your boots on the ground at our test sites to observe machine performance firsthand and build intuitive real-world grounding for how your metrics perform
Take physical ground-truth measurements in the field to validate quantitative metric concepts against actual jobsite results
Build evaluation tooling to triage performance regressions and design agentic feedback loops that accelerate evaluation analysis
Validate metric outputs against domain expert and operator judgment to ensure scoring aligns with real construction standards
Document metric definitions assumptions and failure modes clearly so the team can maintain and extend your scoreboard post-internship
Currently pursuing a BS MS or PhD in computer science robotics statistics operations research or a related quantitative field or bringing equivalent hands-on experience
Strong Python and comfort building data pipelines and analysis tooling
Ability to translate a fuzzy notion of good into a concrete measurable definition and to articulate why that definition is the right one
Solid statistical intuition: understanding variance significance and when a number is actually telling you something
Clear written and verbal communication metric definitions have to be understood and trusted by people who didnt build them
Experience with evaluation frameworks benchmarking or measurement methodology in ML robotics or a related domain
Familiarity with geospatial data point clouds or survey-grade measurement
Prior work with autonomous vehicle or robotics evaluation pipelines
Bedrock Robotics is an Equal Opportunity Employer
Were committed to building a diverse and inclusive workplace. We consider all qualified applicants for employment without regard to race color religion sex sexual orientation gender identity national origin ancestry age disability veteran status genetic information or any other protected characteristic.
Reasonable Accommodations
We want our hiring process to be accessible to everyone. If you need an accommodation to participate in the application or interview process please let your recruiter know so we can support you.
Required Experience:
Intern