Summary
Bolder Apps is a product development studio that partners with startups and established companies to build and scale innovative digital products. The company is seeking a mid-senior LLM Application Engineer to design, ship, and harden production AI pipelines involving structured extraction, classification, evaluation, reliability, and cost optimization. The role collaborates with product, mobile, and QA teams to deliver reliable LLM-powered features for client products.
Responsibilities
- Own production LLM pipelines end to end: ingestion, multimodal model calls, structured records, and storage, including confidence flags, retries, and idempotent rescans
- Design prompt and schema strategies (including schema-aligned or constrained outputs) so results are consistent and product-ready
- Build classification and filtering layers on top of extraction (taxonomy mapping, demographic or audience filters, deduplication, and related cleanup logic)
- Define and run evaluation harnesses (golden sets, regression suites, online metrics) so quality does not regress when prompts, models, or parsers change
- Hit and report against hard quality targets (precision-style gates for completeness, duplicates, incorrect inclusions, image presence, and similar product SLAs)
- Optimize token usage, model tiering, caching, and batching to keep dollars-per-run and latency under control
- Harden reliability for long-running async jobs (timeouts, partial recovery, memory limits, safe production deploys)
- Partner with Flutter / mobile and QA on field contracts, review queues, and incident debugging
- Document architecture and runbooks so ownership is shared, not a single point of failure
- Stay current on Gemini and peer LLM APIs; recommend when to swap models, add fallbacks (e.g. document AI), or tighten schemas
Skills
- We want someone who has shipped LLM apps for real users, not demos
- You should be strong across modern LLMs and especially fluent with Google Gemini (multimodal prompts, structured outputs, failure modes, and cost/latency tradeoffs), with solid experience on other major providers too
- Shipped LLM applications in production (not demos only): prompts, structured outputs, retries, observability, and real failure handling
- Strong hands-on experience with Google Gemini, including multimodal (text + image / document-style) workflows and structured extraction
- Practical experience with at least one other major LLM stack (OpenAI, Anthropic, or similar) and good judgment on when to use which
- Structured extraction from messy inputs: HTML, PDFs, images, and mixed email-like content
- Classification / taxonomy systems on top of LLM outputs
- Evaluation discipline: offline evals, regression suites, and production quality metrics tied to clear acceptance criteria
- Cost and latency awareness: token budgeting, cheaper tiers, caching, batching; can explain dollars-per-run tradeoffs to a PM
- Python backend experience on serverless cloud (Cloud Functions or equivalent) and document stores (e.g. Firestore) or similar GCP patterns
- English at C1 or above for client-adjacent debugging with a PM
- Ownership habits: honest estimates, early blockers, finished releases
- Schema-aligned LLM frameworks (BAML, Instructor, Outlines, or similar)
- Google Document AI or other OCR / document intelligence as a fallback path
- Gmail API / OAuth products and restricted-scope compliance familiarity
- Computer vision for product-image quality checks
- Firebase + GCP ops (secrets, regions, schedulers, cost monitoring)
- Building eval corpora from real production data and iterating until contractual SLAs pass
- Agency or multi-client studio experience
Qualifications
Must Haves
- We want someone who has shipped LLM apps for real users, not demos
- You should be strong across modern LLMs and especially fluent with Google Gemini (multimodal prompts, structured outputs, failure modes, and cost/latency tradeoffs), with solid experience on other major providers too
- Shipped LLM applications in production (not demos only): prompts, structured outputs, retries, observability, and real failure handling
- Strong hands-on experience with Google Gemini, including multimodal (text + image / document-style) workflows and structured extraction
- Practical experience with at least one other major LLM stack (OpenAI, Anthropic, or similar) and good judgment on when to use which
- Structured extraction from messy inputs: HTML, PDFs, images, and mixed email-like content
- Classification / taxonomy systems on top of LLM outputs
- Evaluation discipline: offline evals, regression suites, and production quality metrics tied to clear acceptance criteria
- Cost and latency awareness: token budgeting, cheaper tiers, caching, batching; can explain dollars-per-run tradeoffs to a PM
- Python backend experience on serverless cloud (Cloud Functions or equivalent) and document stores (e.g. Firestore) or similar GCP patterns
- English at C1 or above for client-adjacent debugging with a PM
- Ownership habits: honest estimates, early blockers, finished releases
Nice to Haves
- Schema-aligned LLM frameworks (BAML, Instructor, Outlines, or similar)
- Google Document AI or other OCR / document intelligence as a fallback path
- Gmail API / OAuth products and restricted-scope compliance familiarity
- Computer vision for product-image quality checks
- Firebase + GCP ops (secrets, regions, schedulers, cost monitoring)
- Building eval corpora from real production data and iterating until contractual SLAs pass
- Agency or multi-client studio experience
Benefits
- Fully remote and async-friendly, with required overlap through ~5 PM EST when client or release coordination needs it
- Real autonomy over how you structure prompts, schemas, evals, and deploys. We do not hand you a rigid playbook
- Direct line to PMs, mobile engineers, and decision-makers
- Tooling budget for the LLM and cloud tools you need to move fast
- A peer network of product-minded builders across overlapping client projects
- A team that's always ready to help you grow