06/29/2026 updated


Premium member
100 % availableAI Quality and LLM Evaluation Engineer | Senior SDET | RAG & Agentic Systems Testing
Bucharest, Romania
Only remote
Bachelor in Applied Informatics and Automation, Hyperion University of Bucharest (2015-2019)AI AgentsLLMTesting (Software)Agile TestingAcceptance Test-Driven DevelopmentAtlassian JiraTest AutomationAutomated Testing FrameworkCloud TestingDebuggingLoad TestingData Driven TestsSystem TestingTest DesignTest Execution Engine
AI / LLM Testing
Expertise with DeepEval, RAGAS and Langfuse for LLM evaluation, including traces, prompt versioning, dataset management and evaluation runs. Proficiency in golden dataset regression, faithfulness, answer relevancy and hallucination detection, as well as RAG retrieval metrics such as hit rate, MRR, precision/recall@k and context relevancy. Agentic system testing covering multi-turn state, tool and function calling, and planning loops.
Test Frameworks and Automation Design
Hands-on experience with Playwright (UI, API, Component), Selenium, Cypress and Vitest. Framework and runner design, model based testing with xState, and application of POM and Screenplay patterns to build scalable, maintainable test infrastructure.
CI/CD, Infra and Observability
Practical knowledge of GitLab CI, GitHub Actions, Jenkins, Git, Docker and Kubernetes. Integration of monitoring and observability tools including Datadog, OpenTelemetry and Grafana, with experience in trace driven debugging and connecting evaluation harnesses into CI/CD pipelines as release gates.
Programming Languages
Primary development in Python and TypeScript, used for building production APIs, test frameworks, automation pipelines and evaluation harnesses.
API and Property Based Testing
Experience with Schemathesis, Hypothesis and fast-check for property based testing. Pydantic schema validation and REST, GraphQL, OpenAPI first approaches, with Postman for API exploration and testing.
Test Management and Process
Proficiency with Jira, Xray, TestRail, Qase.io and Allure for test management. Application of shift left and shift right strategies, testing trophy and diamond models, regression planning and release gate decision making.
Model Based Testing
Building model based test systems using xState, enabling generation of large numbers of test scenarios from a single state machine definition and automatic adaptation when product permissions change.
QA Process Design
Experience building QA processes from zero in multiple companies, owning quality strategy within feature squads, reviewing developer-written tests, advising on coverage gaps and test pyramid placement.
Expertise with DeepEval, RAGAS and Langfuse for LLM evaluation, including traces, prompt versioning, dataset management and evaluation runs. Proficiency in golden dataset regression, faithfulness, answer relevancy and hallucination detection, as well as RAG retrieval metrics such as hit rate, MRR, precision/recall@k and context relevancy. Agentic system testing covering multi-turn state, tool and function calling, and planning loops.
Test Frameworks and Automation Design
Hands-on experience with Playwright (UI, API, Component), Selenium, Cypress and Vitest. Framework and runner design, model based testing with xState, and application of POM and Screenplay patterns to build scalable, maintainable test infrastructure.
CI/CD, Infra and Observability
Practical knowledge of GitLab CI, GitHub Actions, Jenkins, Git, Docker and Kubernetes. Integration of monitoring and observability tools including Datadog, OpenTelemetry and Grafana, with experience in trace driven debugging and connecting evaluation harnesses into CI/CD pipelines as release gates.
Programming Languages
Primary development in Python and TypeScript, used for building production APIs, test frameworks, automation pipelines and evaluation harnesses.
API and Property Based Testing
Experience with Schemathesis, Hypothesis and fast-check for property based testing. Pydantic schema validation and REST, GraphQL, OpenAPI first approaches, with Postman for API exploration and testing.
Test Management and Process
Proficiency with Jira, Xray, TestRail, Qase.io and Allure for test management. Application of shift left and shift right strategies, testing trophy and diamond models, regression planning and release gate decision making.
Model Based Testing
Building model based test systems using xState, enabling generation of large numbers of test scenarios from a single state machine definition and automatic adaptation when product permissions change.
QA Process Design
Experience building QA processes from zero in multiple companies, owning quality strategy within feature squads, reviewing developer-written tests, advising on coverage gaps and test pyramid placement.
Languages
EnglishFluentRomanianNative speaker
Project history
Built an internal LLM-powered debugging agent connecting failing test output with Datadog traces, stack traces and recent Jira changes to reduce bug triage time from hours to minutes. Owns quality strategy within the feature squad, maintains and extends the API test framework, applies shift left approach, uses agentic workflows for MCP servers, slash commands and test maintenance, and is responsible for regression planning and release gate decisions in the payments domain.
Built LLM evaluation for a RAG assistant answering researcher questions based on FDA and EMA guidance, separating retrieval quality from generation quality using RAGAS and Langfuse with CI execution. Built a model based testing system for the permissions matrix using xState generating around 1,000 test scenarios. Built a Playwright and TypeScript test platform with POM and tenant/role isolation. Created an AST-based migration tool translating Karate DSL to TypeScript and Playwright. Served as framework SME for onboarding, test pattern reviews and cross-cutting changes.
Built a custom Pytest plugin running dependent API tests in parallel with order resolved by a DAG, reducing runtime while maintaining correctness. Patched and rebuilt the Chromium fork of Qt WebEngine to expose the remote debugging port, enabling replacement of Appium with Selenium through CDP and reducing effort per test from 4 days to 1 day.