Summary
MeridianLink is building a shared AI platform runtime for AI agents used by credit unions and lenders. The AI Engineer will develop trust and explainability capabilities, including multi-agent tracing, evaluation tooling, customer-facing explanations, safety testing, and tenant isolation. The role also collaborates with product engineers and Security Operations while contributing to platform documentation and team development.
Responsibilities
- Completes assigned features and bug fixes independently, with limited need for day-to-day guidance
- Designs components within a well-defined scope; escalates questions on complex system design rather than guessing
- Participates in code review and provides constructive feedback
- Surfaces blockers proactively rather than waiting for check-ins
- Writes tests that cover the functionality they ship
- Monitors and responds to issues with their own work
- Documents decisions and implementation details that others will need later
- Understands how to instrument an agent so a single conversation can be followed from the first user message through every model call, tool call, retrieval, agent handoff, and decision to the final response
- Builds the correlation that ties actions across multiple agents in one workflow into a single readable trace
- Turns raw trace data into an explanation a human can follow: what the agent did, what it relied on, and why it chose that path
- Understands that the explanation an engineer needs and the explanation a borrower or loan officer needs are different, and builds for both
- Works with product engineers on what an agent should disclose at the product surface: what it did, what data it used, how confident it was, and what a human should verify
- Applies judgment about the right level of explanation for a regulated lending and account-opening context
- Familiar with the open-source landscape for LLM observability, tracing, and evaluation, and can evaluate a framework against the platform's needs
- Integrates and extends existing frameworks rather than rebuilding them, and builds what does not exist yet
- Keeps the platform's instrumentation aligned to emerging standards so traces stay portable across tools
- Understands how to measure the quality of non-deterministic output: golden datasets, rubric and LLM-as-judge scoring, and regression baselines
- Builds evaluation checks that run in CI so a model swap, prompt change, or tool change is tested before it reaches customers
- Treats evaluation data as versioned, reviewed code
- Knows the common attack patterns against LLM applications (prompt injection, jailbreaks, tool misuse, data exfiltration) and how to write tests for them
- Understands tenant isolation as a platform guarantee and builds tests that prove one customer's agent can never see another customer's data
- Contributes to the shared guardrail layer so every agent inherits protection without re-implementing it
- Build tracing across the platform's gateway, orchestration, memory, and tool layers, following defined designs from the team lead and architects
- Build the correlation that links actions across agents in a single multi-agent workflow, including handoffs, parallel branches, and retries
- Build the developer-facing trace view that makes a full agent conversation readable to someone who did not write the agent
- Build the explanation layer that turns trace data into a human-readable account of why an agent did what it did
- Build the platform primitives product teams surface to end users: explanation records, confidence and provenance metadata, and "what the agent relied on" summaries
- Partner with product engineers on the Document Request Agent and the MLM agents to land those primitives in real products
- Iterate on the explanation format based on feedback from product teams and customer-facing staff
- Evaluate and integrate open-source observability, tracing, and evaluation frameworks into the platform runtime
- Extend those frameworks where our agent workloads need more than they offer, and contribute fixes and extensions back upstream where it makes sense
- Build the gaps: the pieces of trust and explainability tooling the ecosystem does not provide yet
- Build and maintain components of the platform's evaluation framework: golden dataset management, test runners, scoring pipelines, and regression reporting
- Build the tooling that helps teams create and version golden datasets from de-identified real traffic and synthetic cases
- Run model and prompt comparisons for the platform's shared components and report what changed
- Build red-team and adversarial test suites covering prompt injection, jailbreak attempts, tool misuse, and data exfiltration, and run them in the platform's release process
- Build the automated test suite that proves agents on the shared runtime cannot cross tenant boundaries through memory, retrieval, tool calls, or model context
- Share red-team findings and new attack patterns with Security Operations and pull their threat models into the platform's tests
- Participate in design discussions and code reviews; give and receive feedback constructively
- Support onboarding of L1 AI Engineer teammates; share context and help them get unblocked
- Contribute to documentation that reduces tribal knowledge on the team
Skills
- 3+ years of professional software engineering experience, delivering features independently in a production environment
- Solid understanding of algorithms, data structures, and software design fundamentals
- Proficiency with standard development tooling: Git, Docker, automated testing, and modern scripting languages
- Active daily use of AI-assisted development tools
- Bachelor's degree in Computer Science, Software Engineering, or equivalent experience
- Hands-on experience building software that integrates large language models (LLM APIs, agent frameworks, RAG pipelines, or similar), in production or in substantial personal or open-source projects
- Proficiency in Python with production experience
- Hands-on experience in at least one of the following, with real interest in growing into the others: LLM observability and tracing, LLM evaluation and testing, agent frameworks and multi-agent orchestration, or application security testing
- Experience with distributed tracing or observability tooling in a production system
- Strong automated testing instincts, including an interest in how to test systems that do not return the same answer twice
- Experience running workloads on Azure or AWS, including the basics of identity and access management, networking, and secrets management
- TypeScript a plus
- Experience with OpenTelemetry, including the GenAI semantic conventions, or OpenLLMetry
- Experience with LLM observability and evaluation tools (Langfuse, Arize Phoenix, LangSmith, Braintrust, promptfoo, DeepEval, or equivalent)
- Contributions to open-source AI observability, evaluation, or agent framework projects
- Experience with cloud-managed model services such as AWS Bedrock or Azure OpenAI
- Experience building or operating multi-tenant SaaS systems where tenant isolation was a hard requirement
- Prior experience in financial services, fintech, or another regulated industry where explainability shaped technical decisions
- Experience building developer-facing debugging or visualization tools
Qualifications
Must Haves
- 3+ years of professional software engineering experience, delivering features independently in a production environment
- Solid understanding of algorithms, data structures, and software design fundamentals
- Proficiency with standard development tooling: Git, Docker, automated testing, and modern scripting languages
- Active daily use of AI-assisted development tools
- Bachelor's degree in Computer Science, Software Engineering, or equivalent experience
- Hands-on experience building software that integrates large language models (LLM APIs, agent frameworks, RAG pipelines, or similar), in production or in substantial personal or open-source projects
- Proficiency in Python with production experience
- Hands-on experience in at least one of the following, with real interest in growing into the others: LLM observability and tracing, LLM evaluation and testing, agent frameworks and multi-agent orchestration, or application security testing
- Experience with distributed tracing or observability tooling in a production system
- Strong automated testing instincts, including an interest in how to test systems that do not return the same answer twice
- Experience running workloads on Azure or AWS, including the basics of identity and access management, networking, and secrets management
Nice to Haves
- TypeScript a plus
- Experience with OpenTelemetry, including the GenAI semantic conventions, or OpenLLMetry
- Experience with LLM observability and evaluation tools (Langfuse, Arize Phoenix, LangSmith, Braintrust, promptfoo, DeepEval, or equivalent)
- Contributions to open-source AI observability, evaluation, or agent framework projects
- Experience with cloud-managed model services such as AWS Bedrock or Azure OpenAI
- Experience building or operating multi-tenant SaaS systems where tenant isolation was a hard requirement
- Prior experience in financial services, fintech, or another regulated industry where explainability shaped technical decisions
- Experience building developer-facing debugging or visualization tools
Benefits