Politics Content Reviewer (AI Evaluation)

Gramian Consulting Group Bangladesh
Apply Now

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

Role Overview We are looking for an experienced Political Science Domain Expert to contribute to the evaluation and improvement of Large Language Models. The role involves creating advanced prompts, assessing AI-generated responses, identifying knowledge gaps, and developing high-quality evaluation data. The work may cover political science, governance, elections, public policy, political institutions, comparative politics, and international relations. The successful candidate will apply strong domain knowledge and reliable sources to evaluate factual accuracy, reasoning quality, completeness, and nuance.

CONTRACT: Contractor assignment, 8 weeks COMMITMENT: Full-time, 40 hours per week with at least 4 hours of PST overlap LOCATIONS: Remote - Bangladesh, Brazil ,Colombia, Egypt, Ghana, India, Indonesia, Pakistan, Turkey, Vietnam PROCESS: Initial screening followed by domain manager review

Key Responsibilities • Create advanced political science prompts across a range of topics and difficulty levels. • Evaluate AI-generated responses for factual accuracy, reasoning quality, completeness, and nuance. • Review content related to governance, elections, public policy, political institutions, and international relations. • Identify hallucinations, outdated claims, logical inconsistencies, bias, and complex edge cases. • Develop benchmark datasets and adversarial test cases for model evaluation. • Verify conclusions using reliable, relevant, and up-to-date references. • Provide clear, objective, and evidence-based feedback on model outputs. • Maintain accurate documentation and consistent annotation quality. • Collaborate with AI researchers and project teams to improve model performance.