AutoResearch + PromptFoo = AutoPrompter. Run it with Neo AI Engineer
-
Updated
Apr 16, 2026 - Python
AutoResearch + PromptFoo = AutoPrompter. Run it with Neo AI Engineer
Your guide to the Agentic AI evolution. **Prompting Blueprints** offers a curated collection of concepts and tactics for building autonomous AI workflows. Master tool-specific playbooks, backed by structured prompt packs and rigorous evaluations for the latest AI models.
This repository contains the code snippets used in "LLM Prompt Engineering For Developers"
prompt-evaluator is an open-source toolkit for evaluating, testing, and comparing LLM prompts. It provides a GUI-driven workflow for running prompt tests, tracking token usage, visualizing results, and ensuring reliability across models like OpenAI, Claude, and Gemini.
Community Plugin for Genkit to use Promptfoo
面向开发者的中文 Eval Harness 源码教材:解析 lm-evaluation-harness、Inspect AI、OpenAI Evals、Promptfoo、DeepEval 与 Harbor,覆盖任务、运行、评分、统计与发布门禁。
Orchestrator/sub-agent RCA system with a production-grade eval and observability harness — LangGraph agents, Ragas/Promptfoo evals, OTel tracing.
Agent 测试转型手册:资深自动化测试工程师的 AI QE 学习路径 + 练手代码
VCL VibeBench — open frontier-model comparisons you can fork and re-run. By Vibe Coder's Life.
Compare models across prompts, test domains, and scenarios.
Python + SQL toolkit for LLM threat intelligence — flags prompt injection, jailbreaks, and data exfiltration mapped to OWASP LLM Top 10 (2025) & MITRE ATLAS.
Sample project demonstrates how to use Promptfoo, a test framework for evaluating the output of generative AI models
AI Testing Framework using Promptfoo for automated LLM evaluation, red teaming, and CI/CD integration with Jenkins. Supports prompt validation, security testing, custom assertions, API-based model testing, and quality gates for enterprise GenAI applications.
Quickstart guide for using PromptFoo to evaluate LLM prompts via CLI or Colab.
Real DuckDB Quack infrastructure for multi-agent Werewolf: containerized player nodes, Quack gateway federation, browser runner, and local/hosted LLM evals.
Prompt optimization and benchmarking for local LLMs: picks the right prompt technique for your task and measures it on Ollama. DSPy-compatible.
To associate your repository with the promptfoo topic, visit your repo's landing page and select "manage topics."