Object logo
Object
Posted 54 days agoVerified live 1d ago

Machine Learning Engineer Intern — AI for Science

Brief overview

Remote
UndergradOr in progress
Sponsors visasStated in posting
PythonPandasNumPyMachine LearningSQLDatabase Schema DesignData IntegrationSchema MatchingEntity ResolutionLarge Language Models (LLMs)Retrieval-Augmented Generation (RAG)LangChain

About the company

Deep technology is the rocket that can carry human civilization to its next frontier — yet building this rocket remains painfully slow, fragile, and error-prone, held back by fragmented data, legacy manual workflows, and a scarcity of domain talent.

Job description

Summary

Object Tech, Inc. is a technology company providing full-stack AI solutions for technology R&D and production, combining expertise in scientific and engineering disciplines. The Machine Learning Engineer Intern will help build an AI agent for analyzing historical scientific data by working across machine learning, large language models, and data engineering. The role involves profiling and integrating heterogeneous datasets, developing LLM-based table relationship inference, evaluating system performance, and collaborating with domain experts.

Responsibilities

  • Exploring and profiling messy, real-world scientific tables (CSV/Excel), extracting structure and metadata
  • Contributing to a relationship-inference pipeline that combines classical data-integration signals (schema and value matching, key detection) with LLM-based semantic reasoning over column meanings and provenance
  • Supporting the development of evaluation methods that measure system performance against expert-provided ground truth
  • Load, clean, and normalize heterogeneous tabular datasets and build reusable data-profiling tooling
  • Prototype and iterate on an LLM/agent pipeline that classifies how pairs of tables relate (parallel / shared / hierarchical)
  • Design and maintain benchmarks and metrics to evaluate system accuracy, and run experiments to improve performance
  • Work with database schemas, including joins, keys, and data modeling, to represent and query discovered relationships
  • Collaborate with mentors and domain scientists to understand requirements and turn feedback into concrete improvements
  • Document findings, experiments, and technical approaches throughout the project

Skills

  • BS/MS (in progress or completed) in Data Science, Computer Science, or a closely related field
  • Strong Python, including data-wrangling libraries (e.g., pandas, NumPy)
  • Solid grounding in machine learning fundamentals
  • Hands-on database experience, including SQL, schema design, and joins
  • Experience with LLMs / AI agents (prompting, RAG, tools like LangChain)
  • Prior work with scientific or experimental datasets
  • Familiarity with data integration, schema matching, or entity resolution
  • Good software habits such as version control, testing, and clear documentation

Qualifications

Must Haves

  • BS/MS (in progress or completed) in Data Science, Computer Science, or a closely related field
  • Strong Python, including data-wrangling libraries (e.g., pandas, NumPy)
  • Solid grounding in machine learning fundamentals
  • Hands-on database experience, including SQL, schema design, and joins

Nice to Haves

  • Experience with LLMs / AI agents (prompting, RAG, tools like LangChain)
  • Prior work with scientific or experimental datasets
  • Familiarity with data integration, schema matching, or entity resolution
  • Good software habits such as version control, testing, and clear documentation

Benefits

  • Fully remote with flexible schedule
  • Collaborate with team members from leading tech firms (including MAMAA)
  • Work on high-impact, real-world AI marketing projects
  • Great for students (supports CPT/OPT)
  • Opportunity for recommendation letters, referrals, and future growth
  • Direct mentorship from professionals in product, marketing, and AI

More jobs like this