Senior QA / ML Tester w Katowice, Polska is listed on Jobeax. Browse 110,000+ vacancies available.
We are looking for a Senior QA / ML Tester to join the AI Platform team and take ownership of quality assurance for the Agent Evaluation Framework built on AWS AgentCore. This role involves designing, implementing, and maintaining a functional test suite that validates the correctness of AI agent evaluation pipelines — covering on-demand integration testing, online sampling accuracy, and multi-evaluator execution — and delivering a feasibility assessment for non-AgentCore runtime evaluation scenarios. This is a hands-on, production-focused role at the intersection of software quality engineering and AI/ML system testing, operating within an Agile delivery team and contributing to the reliability of enterprise-grade agentic AI infrastructure. Responsibilities Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good / known-bad session pair validation, multi-evaluator execution correctness, and edge case handlingDevelop and maintain integration tests for on-demand evaluation mode, integrated into the CI/CD pipeline with automated execution on each buildValidate online mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidenceConduct and document a feasibility assessment for non-AgentCore runtime evaluation: analyze alternative runtimes, define evaluation methodology, and deliver a structured findings reportTest OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrityCollaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gatesMaintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence/JiraParticipate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectivesContribute to EngX practices such as code review of test scripts, CI/CD pipeline integration, and test coverage reporting Requirements 5+ years of production experience in QA automation or ML/AI system testingProven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validationProficiency in Python test automation, including pytest, fixtures, parametrize, and mocking (https://jobeax.com/link/Vhs9R0o6sLJV4mN3, moto)Knowledge of AWS AgentCore Evaluation API, covering on-demand and online evaluation modesFamiliarity with OpenTelemetry / ADOT for trace-based evaluation input testing and trace structure validationSkills in REST API testing, including request/response validation and authentication (SigV4, bearer tokens)Experience with CI/CD integration using GitHub Actions, Jenkins, or equivalent, including test pipeline configurationBackground in test data management, including known-good / known-bad session pair design and synthetic trace generationExpertise in functional and integration test design for AI/ML evaluation pipelinesCompetency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysisUnderstanding of LLM/agent evaluation concepts, such as correctness scoring, sampling strategies, and evaluator chainingAbility to work independently after onboarding, manage own tasks, report status, and escalate blockers proactivelyStrong analytical skills to define test scenarios from ambiguous or evolving specificationsEnglish B2+ level, written and verbal, for daily collaboration with distributed teams Nice to have Hands-on experience with AWS AgentCore Evaluation API or AWS Bedrock testingExperience testing OpenTelemetry / distributed tracing pipelinesFamiliarity with multi-evaluator execution patterns and correctness validation strategiesExperience writing feasibility assessments or technical reports for stakeholdersKnowledge of agentic AI frameworks (LangGraph, Strands Agents) sufficient to understand evaluation contractsISTQB CT-AI certification or equivalent AI testing qualificationExperience with AI Ready / AI Practitioner practices at EPAM (prompt engineering, AI-assisted test design) We offer We gather like-minded people:Top tech minds driving innovation in AI, cloud and digital platform modernizationSupportive team and agile, startup-like cultureHybrid by design mode and opportunity to work remotely within PolandChance to work abroad for up to 60 days annuallyBusiness-driven relocation opportunitiesWe provide growth opportunities:Career development programsThought leadership, mentoring, soft skills and well-being programsCertification (Anthropic, Gemini, GCP, Azure, AWS)English classesWe cover it all:Stable payParticipation in the Employee Stock Purchase Plan with a 15% discountBenefits package (health insurance, multisport, shopping vouchers)Referral bonuses up to $2,000Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and moreCorporate, social and well-being eventsPlease, note:Benefits listed above are available to employees onlyWe are open for working with Contractors. Terms of B2B cooperation agreements are agreed individuallyWe will reach out to selected candidates exclusively EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.