inhousefyi
← Back to listings

LLM / Agentic Evaluation Rig Engineer

PhizenixHyderabad, India · Posted 7 days ago
Full-time
Apply now

Description

We are looking for an LLM / Agentic Evaluation Rig Engineer to build the system that decides whether our AI output is good enough to ship. Because our commentary sits next to externally reported financials, we cannot rely on vibes — grounding, faithfulness, and hallucination have to be measured, tracked, and gated before anything reaches a customer.

You own the evaluation infrastructure: the datasets, the scorers, the harnesses, and the CI gates that hold the AI and agentic layers to a hard quality bar. You are the team's source of truth on whether a model, prompt, or agent change is actually an improvement — and the one who blocks it if it isn't.

What makes this role different

  • You define "good enough to ship" — your gates block regressions in grounding and faithfulness from reaching production.
  • Evidence over vibes — every claim is checked against the verified source data it must be grounded in.
  • Agentic evaluation — you evaluate multi-step reasoning flows, not just single prompts.
  • Real leverage — your rig is how the whole AI team moves fast without breaking trust.

Responsibilities

Datasets & Scorers (35%)

  • Build and curate evaluation datasets, including adversarial and edge-case sets with ground-truth labels
  • Build scorers for grounding, faithfulness, hallucination, factual consistency, and structured-output validity
  • Combine rule-based checks, reference-based metrics, and LLM-as-judge where appropriate
  • Verify generated claims map to verified source data — no unsupported statements

Harnesses & CI Gates (30%)

  • Build harnesses that run evaluations reproducibly across model, prompt, and agent versions
  • Wire evaluation into CI so grounding / faithfulness regressions block releases
  • Track quality over time with dashboards and clear pass / fail thresholds

Agentic Evaluation (25%)

  • Evaluate multi-step / agentic flows — routing, tool-use, verification, confirmation
  • Build trace capture and step-level scoring for agent runs
  • Detect where a flow silently degrades

Collaboration (10%)

  • Partner with the Staff AI Engineer to turn findings into model / prompt / orchestration improvements
  • Partner with QA to integrate AI evaluation into the broader release process

Technical Stack

Evaluation

  • LLM eval frameworks (promptfoo, DeepEval, Ragas, LangSmith)
  • LLM-as-judge, reference-based metrics
  • Dataset / ground-truth curation

AI & Orchestration

  • LLM APIs & managed LLMs (Bedrock / Vertex / Azure OpenAI)
  • RAG & agentic patterns (LangGraph)
  • Structured-output validation

Engineering

  • Python
  • CI/CD (GitHub Actions)
  • Dashboards & metrics tracking

What You'll Build in Year One

  • A labeled evaluation dataset suite (including adversarial cases) for the generation and agentic layers.
  • A scorer library for grounding, faithfulness, hallucination, and structured-output validity.
  • A reproducible harness wired into CI that blocks releases on quality regressions.
  • Step-level trace capture and scoring for agentic flows, with dashboards leadership can trust.

Required Qualifications

Core

  • 4+ years in software / ML engineering, with hands-on work building LLM evaluation or quality tooling.
  • Real understanding of grounding, faithfulness, and hallucination — and how to measure them rigorously.

Technical

  • Strong Python and solid engineering practices (reproducibility, CI/CD).
  • Comfort designing evaluation for non-deterministic systems without producing flaky or meaningless metrics.
  • Familiarity with LLM eval frameworks and LLM-as-judge patterns.

Nice-to-Have

  • Experience evaluating agentic / multi-step LLM systems.
  • Familiarity with RAG, structured output, and managed LLMs in-VPC.
  • FinTech / financial-services domain or other high-stakes, correctness-critical AI.
  • Background in statistics or measurement / metrics design.

Similar jobs

BackbaseVietnam

The Job in short The Data Enablement team is here to enable every team in the organisation with their data needs. Our job starts the moment that data enters our platform and ends when it reaches whoever needs it. We are…

Full-time
BrightAI CorporationPalo Alto, California, United States

Est. 200,000 USD

Sr. AI Engineer – LLM, RAG BrightAI is a high-growth Physical AI company transforming how businesses interact with the physical world through intelligent automation. Our AI platform processes visual, spatial, and tempora…

Full-time
WorkatoPalo Alto, California, United States

Est. 168,000 USD

About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…

Full-time
AirPittsburgh, Pennsylvania, United States

Est. 124,000 USD

Company Description Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. Today, a dangerous Readiness Gap exists between what the front…

Full-time
N-iXRemote

We're looking for an engineer with hands-on experience building and evaluating GenAI services - from RAG and agentic reasoning systems to production-grade LLM deployments. You'll work closely with Frontend and Backend te…

Full-timeRemote
10a LabsRemote

Est. 135,000 USD

About 10a Labs: 10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. Our adversarial red teaming, model evaluations…

Full-timeRemote
Staff AI Engineer2 months ago
WorkatoRemote

About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…

Full-timeRemote

Est. 140,000 USD

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence…

Full-timeRemote
AvePointSingapore, Singapore

We are looking for a highly skilled AI Engineer specializing in Large Language Models (LLMs) and Agentic AI. You will architect, build, and deploy production-grade LLM applications — from intelligent knowledge bases and…

Full-time

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence…

Full-timeRemote
GradialSeattle, Washington, United States

Est. 175,000 USD

Gradial is the marketing operations system of work that helps marketers and creatives move from idea to execution faster. Our platform orchestrates across martech stacks, workflows, and people to automate marketing execu…

Full-time
CodeRoadRemote

Est. 120,000 USD

About CodeRoad CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our…

Full-timeRemote
KayzenRemote

Remote from EMEA or Bangalore Hello 👋 I am Servesh, Co- founder and CTO at Kayzen, and I am now looking for a LLM Platform Engineer/Lead who will be part of our Engineering team. 🙌 But wait, you have not heard of Kayze…

Full-timeRemote
Zone 5 TechnologiesUnited States

Est. 141,000 USD

At Zone 5 Technologies, we're redefining what's possible in unmanned aircraft systems. Our team of engineers and innovators is developing cutting-edge autonomous solutions that push the boundaries of UAS technology - sol…

Full-time
MitratechRemote

Est. 120,000 EUR

At Mitratech, we are a team of technocrats focused on building world-class products that simplify operations in the Legal, Risk, Compliance, and HR functions. We are a close-knit, globally dispersed team that thrives in…

Full-timeRemote
CodeRoadRemote

Est. 120,000 USD

About CodeRoad CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our…

Full-timeRemote
MeridialRemote

Est. 120,000 USD

Are you an AI QA expert eager to shape the future of AI? Large-scale language models are evolving from clever chatbots into enterprise-grade platforms. With rigorous evaluation data, tomorrow’s AI can democratize world-c…

Full-timeRemote

Job Overview: We are looking for a Senior GenAI Developer to design, build, and productionize agentic AI systems—LLM-powered agents that can plan, use tools, orchestrate workflows, and operate reliably under enterprise c…

Full-time
WorkatoSan Francisco, California, United States

Est. 80,000 USD

About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…

Full-time
WorkatoSofia, Bulgaria

About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…

Full-time
R/GA Careers PageLondon, United Kingdom

Est. 80,000 GBP

About R/GA R/GA is an independent creative innovation company built for the intelligence age. We harness the power of design and technology to create more valuable experiences for people and brands. From architecting ada…

Full-time
ArtefactNew York, New York, United States

Est. 200,000 USD

About the job Do you think like a management consultant, thrive in a startup environment, and can’t stop thinking about the intersection of data, technology, and marketing? With over 2000 employees, offices on five conti…

Full-time
CoupangBeijing, China

Senior Staff/Staff Machine Learning Engineer-BJ/SH Company Introduction We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out…

Full-time
AI Researcher3 months ago
WorkatoSingapore, Singapore

About Workato Workato delivers enterprise infrastructure for the agentic era, redefining iPaaS and helping enterprises unify data, applications, processes, and AI into a single, governed platform. A leader in Enterprise…

Full-time
Lead AI Engineer7 months ago
AirPittsburgh, Pennsylvania, United States

Est. 141,000 USD

Company Description Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. Today, a dangerous Readiness Gap exists between what the front…

Full-time
ArionkoderRemote

Est. 120,000 USD

Arionkoder is an AI consulting firm helping companies build, embed, and own AI systems that drive real business outcomes. We combine Product Development, Artificial Intelligence, and Team Augmentation to craft digital pr…

Full-timeRemote

Est. 205,500 USD

Principal AI Engineering Architect We're looking for a Principal AI Engineering Architect to lead the design and delivery of complex, multi-domain systems spanning cloud, data, and AI — with deep, hands-on mastery of mul…

Full-timeRemote
LinkedIn Job WrappingNew York, New York, United States

Est. 200,000 USD

About the job Do you think like a management consultant, thrive in a startup environment, and can’t stop thinking about the intersection of data, technology, and marketing? With over 2000 employees, offices on five conti…

Full-time
LLMOps Engineer25 days ago
ALX AfricaRemote

About ALX Africa ALX Africa, a non-profit organisation under the ALX Foundation, is dedicated to unlocking the potential of Africa's digital future. Formerly part of Sand Tech Holdings, we've embarked on an independent j…

Full-timeRemote
SQUADRemote

Team Summary Our distributed team is looking for an experienced Applied Scientist with a strong background in Large Language models to develop high-performance Generative AI features across Cloud and Edge environments. J…

Full-timeRemote