QA Lead- AI Evaluation & Quality
NeGD is currently inviting applications for the position of QA Lead – AI Evaluation & Quality on a contractual basis.
| Position | QA Lead – AI Evaluation & Quality |
| No. of Positions | 01 |
| Last Date | 30th September 2026 |
Responsibilities
Evaluation Strategy, Standards & Quality Gates
- Own the programme’s AI evaluation methodology and make evaluation a formal gate in the delivery lifecycle — defining, publishing and enforcing common metrics, acceptance thresholds and quality gates
- Select the appropriate evaluation approach per use case — offline benchmarks, human evaluation, LLM-as-judge, or online/production evaluation — and document the rationale
- Ensure Responsible AI evaluation is built into the standard delivery flow rather than added afterwards, aligned to the MeitY Responsible AI advisory and the IndiaAI Safe & Trusted AI framework
Evaluation Infrastructure & Reusable Harnesses
- Design and build automated evaluation harnesses, benchmark and golden datasets, and regression suites that pods reuse across ministry deployments — mirroring the programme’s build-once, reuse-many model
- Wire evaluations into CI/CD as release gates; version datasets, prompts and evaluation configurations so results are reproducible and audit-ready
- Publish reusable evaluation sets, rubrics, harnesses and templates to AIKosh under standard metadata (and to OpenForge where code is shared) for national reuse
Model & Output Quality Evaluation
- Evaluate models with task-appropriate quantitative metrics — precision/recall/F1, ROC-AUC/PR-AUC, calibration, MAE/RMSE and task-specific measures — reported with statistical rigour, including uncertainty, adequate sample sizes and vigilance against metric gaming
- Evaluate AI outputs against defined quality dimensions — accuracy, groundedness, reliability, consistency, relevance and adherence to business and policy requirements — rather than pass/fail alone
Gen-AI, LLM & RAG Evaluation
- Assess LLM and RAG systems for hallucination, factual consistency, groundedness and citation faithfulness, relevance, prompt robustness, input validation and non-deterministic behaviour
- Measure retrieval quality and generation quality separately in RAG pipelines using RAGAS, DeepEval, TruLens or LangSmith
- Use LLM-as-judge where appropriate, calibrated against human labels, controlling for known judge biases (position, verbosity, self-preference), and defaulting to human evaluation where automated judgement is unreliable
Agentic AI Evaluation
- Evaluate agents on their trajectories, not only final answers — tool selection and tool-call correctness, multi-step reasoning and workflow completion, agent handoffs, failure recovery, memory behaviour and end-to-end task-completion accuracy
- Validate human-in-the-loop controls — review, escalation, override, feedback and exception handling — wherever AI decisions require human supervision
Safety, Adversarial & Responsible AI Testing
- Under the technical guidance of the AI Safety Researcher, operationalise adversarial and red-team findings — jailbreaks, prompt injection, indirect prompt injection, unsafe behaviour and sensitive-data (PII) leakage — into standard, reusable test suites the pods apply, and verify guardrail effectiveness (Llama Guard, NeMo Guardrails, Guardrails AI or equivalent)
- Evaluate fairness by measuring performance across demographic, linguistic, geographic and other relevant cohorts using appropriate fairness metrics (Fairlearn, AI Fairness 360)
- Author the QA-side safety and quality evidence — model cards, dataset sheets, evaluation logs and Responsible AI records — for approval
Data & Dataset Quality Validation
- Validate training, evaluation and inference data for missing values, duplicates, schema conformance, class imbalance, data leakage, distribution/covariate shift, feature consistency and train–serve skew, using tools such as Great Expectations, Evidently or Deepchecks
- Curate high-quality evaluation datasets that reflect real citizen usage, edge cases and Indian-language and regional diversity
Production Monitoring & Quality Operations
- Monitor live AI applications for quality and retrieval degradation, drift (data, concept and model), accuracy and latency regressions, token/cost behaviour, failures and safety/compliance violations
- Define alerts, thresholds and dashboards for model-quality, latency, error-rate, drift and compliance signals; support quality-incident response and feed production findings back into the evaluation sets
Human Evaluation Operations
- Design evaluation rubrics and run structured human evaluation and annotation; manage annotator quality and inter-annotator agreement; build preference/label datasets that ground and validate the automated metrics
Cross-Pod Leadership, Mentoring & Governance
- Provide technical direction to the AI QA Engineers in the team; set shared standards and harnesses, keep pod-level practice consistent and current, and drive adoption of sound, current evaluation techniques while screening out hype
- Turn quality evidence into decision-ready go/no-go recommendations for deployments, and represent AI quality in programme governance and monthly reviews
- Contribute AI quality and evaluation requirements to NeGD RFQs and RDRs, and review agency deliverables for evaluation and quality posture
Important Links
| Download Detailed Notification | Click Here |
| Apply Here | Click Here |
| Official Website | Click Here |
About National e-Governance Division (NeGD)
The National e-Governance Division (NeGD) is an independent business division under the Digital India Corporation, Ministry of Electronics and Information Technology. NeGD has been playing a pivotal role in supporting MeitY in Programme Management and implementation of e-Governance projects and initiatives undertaken by various Ministries/ Departments, both at the Central and State levels.
NeGD has been spearheading several innovative initiatives under the aegis of the Digital India Programme. Those have been developed keeping the vision areas of Digital India at the core- providing digital infrastructure as a core utility to every citizen, governance and services on demand and in particular, digital empowerment of the citizens of our country; some of these initiatives include DigiLocker, UMANG, Poshan Tracker, OpenForge Platform, API Setu, National Academic Depository, Academic Bank of Credits, Learning Management System.