Prompt Junky

Evaluation & Testing

Frameworks and platforms for scoring prompt and agent output: eval harnesses, LLM-as-judge tooling, benchmarks, red teaming and observability.

97 prompt tools / page 2 of 3

Langtrace

Prompt Managers & Versioning

An open-source, OpenTelemetry-based observability tool for LLM applications that captures traces, manages prompt versions in a registry and runs evaluations on collected data.

langtrace.ai

MLflow

Prompt Managers & Versioning

The widely used open-source ML platform, which now includes a prompt registry for versioning prompt templates alongside tracing and LLM evaluation for generative applications.

mlflow.org

Prompt flow

Prompt Managers & Versioning

A Microsoft toolkit for building LLM applications as directed flows of prompts, Python code and tools, with batch runs, built-in evaluation flows and a VS Code visual editor.

microsoft.github.io

Chainlit

Evaluation & Testing

An open framework and hosted service for building conversational AI interfaces with built-in observability, feedback collection and evaluation of the conversations produced.

chainlit.io

Langfuse

Prompt Managers & Versioning

An open-source LLM engineering platform combining tracing, evaluations and a prompt management layer where prompts are versioned, labelled and fetched at runtime by the SDK.

langfuse.com

FastChat

Evaluation & Testing

The LMSYS platform for serving and evaluating chat models, including the MT-Bench multi-turn judge harness and the infrastructure used for crowd-sourced model comparisons.

github.com

WhyLabs

Evaluation & Testing

An observability platform for AI and data pipelines that profiles inputs and outputs to detect drift and quality problems, including text from language model applications.

whylabs.ai

Arize AX

Evaluation & Testing

Arize's commercial platform for agent observability and evaluation, adding hosted tracing, online evaluators and issue tracking on top of the open-source Phoenix project.

arize.com

Superagent

Evaluation & Testing

An open-source guardrail layer that inspects model traffic for prompt injection, data leakage and harmful output, and can be embedded in an application or run as a proxy.

superagent.sh

Weights & Biases

Evaluation & Testing

The experiment tracking company whose Weave product records LLM calls, versions prompts and datasets as objects, and runs scored evaluations for generative applications.

wandb.ai

LightEval

Evaluation & Testing

Hugging Face's evaluation toolkit for running benchmark tasks across several inference backends, with custom task and metric definitions and detailed per-sample output.

huggingface.co

Prometheus Eval

Evaluation & Testing

An open-source evaluator language model and toolkit for grading responses against a user-supplied scoring rubric, offered as an alternative to proprietary judge models.

github.com

Pydantic Logfire

Prompt Managers & Versioning

An OpenTelemetry-based observability service from the Pydantic team with first-class instrumentation for LLM calls and agent runs alongside ordinary application traces.

pydantic.dev

OpenEvals

Evaluation & Testing

A LangChain package of ready-made evaluators for common LLM checks such as correctness, conciseness and structured-output matching, usable inside or outside LangSmith.

github.com

OpenLLMetry

Prompt Managers & Versioning

A set of OpenTelemetry instrumentations for LLM frameworks and providers from Traceloop, exporting prompt and completion spans to any compatible observability backend.

traceloop.com

Purple Llama

Evaluation & Testing

Meta's collection of trust and safety tools for language models, including the CyberSecEval benchmark suite and the Llama Guard family of input and output classifiers.

github.com

OpenLIT

Evaluation & Testing

An OpenTelemetry-native observability platform for LLM and agent applications that also includes a prompt hub for versioned prompts and a vault for provider secrets.

docs.openlit.io

W&B Weave

Prompt Managers & Versioning

The Weights & Biases toolkit for LLM applications, automatically versioning prompts, code and datasets as objects and tracking every call for comparison and scoring.

wandb.me

Latitude

Prompt Managers & Versioning

An open-source platform where prompts are written in a templating syntax, versioned in projects, evaluated against datasets and then deployed as callable endpoints.

latitude.so

Datadog LLM Observability

Evaluation & Testing

Datadog's product for tracing agent and LLM calls inside existing application monitoring, with offline experimentation and quality evaluations on production spans.

datadoghq.com

Judgeval

Evaluation & Testing

An open-source stack for tracing agents and running evaluations on those traces, including custom judge models and regression tests triggered from production data.

judgmentlabs.ai

Lakera

Evaluation & Testing

An AI security company whose products test and defend applications against prompt injection and jailbreaks, including a public game used to collect attack prompts.

lakera.ai

PromptTools

Evaluation & Testing

An open-source toolkit for experimenting with prompts across models and vector databases, running grids of prompt variants from a notebook and scoring the results.

prompttools.readthedocs.io

Confident AI

Evaluation & Testing

The hosted platform built around the open-source DeepEval framework, adding shared datasets, regression comparison between prompt versions and production tracing.

confident-ai.com

Klu

Prompt Managers & Versioning

A platform for designing, deploying and optimizing LLM applications, with collaborative prompt editing, versioned deployments and evaluation of saved test cases.

klu.ai

LangSmith

Prompt Managers & Versioning

LangChain's hosted platform for tracing, evaluating and managing prompts, including a prompt hub with versioned prompts that can be pulled by the LangChain SDKs.

langchain.com

promptmap

Evaluation & Testing

A scanner that automatically tests a custom LLM application's system prompt against prompt injection and jailbreak rules, reporting the attacks that got through.

github.com

Arthur

Evaluation & Testing

An AI governance and monitoring platform that gives security, risk and engineering teams a shared control plane for evaluating and supervising deployed models.

arthur.ai

Guardrails AI

Evaluation & Testing

A Python framework that wraps model calls with input and output validators from a shared hub, re-asking the model or failing closed when a check does not pass.

guardrailsai.com

LangWatch

Prompt Managers & Versioning

An open-source platform for LLM evaluation and agent testing with prompt versioning, scenario-based agent simulations and an optimization studio built on DSPy.

langwatch.ai

Adaline

Prompt Managers & Versioning

An observability and evaluation platform that turns production traces into datasets and prompt improvements, marketed for teams running self-improving agents.

adaline.ai

Agentic Security

Evaluation & Testing

An open-source red-teaming kit that fuzzes LLM endpoints with adversarial prompt datasets and multimodal attacks, producing a report of which defences failed.

agentic-security.vercel.app

LMMs-Eval

Evaluation & Testing

An evaluation toolkit for multimodal models covering text, image, video and audio tasks with a unified interface, used to benchmark vision-language prompting.

lmms-lab.com

Patronus AI

Evaluation & Testing

An evaluation platform offering managed judge models and test suites for scoring model output, including document-heavy workflows measured by pages processed.

patronus.ai

Prompt Security

Evaluation & Testing

A security product, now part of SentinelOne, that inspects prompts and responses across an organisation to block injection attacks and sensitive data leakage.

prompt.security

Respan (Keywords AI)

Prompt Managers & Versioning

An LLM engineering platform, previously branded Keywords AI, that combines observability, evaluations, prompt optimization and a model gateway behind one API.

keywordsai.co