-
Notifications
You must be signed in to change notification settings - Fork 0
Home
π§ Wiki is work in Progress
The NiriZan Wiki is currently being built. Pages, explanations, and navigation will be added and refined as the project evolves.
NiriZan is an open-source continuous evaluation infrastructure and Python framework for production AI systems. It is designed to continuously evaluate LLMs, RAG applications, AI agents, and other probabilistic AI systems throughout their lifecycle.
NiriZan combines AI evaluation, LLM-as-a-Judge reliability, judge drift attribution, fixed anchor sets, repeatable rescoring, statistical regression detection, statistical gating, trust-weighted health scoring, and CI/CD-integrated regression gates into a single architecturally disciplined framework.
Inspection through Measurement.
This Wiki is intended to become NiriZan's knowledge base: a place for concepts, architectural explanations, statistical methodology, design rationale, practical guides, research directions, and answers to questions about the project.
AI systems are probabilistic. The same system can produce different outputs over time, and changes in models, prompts, retrieval systems, datasets, infrastructure, or evaluation judges can affect observed quality.
Traditional software testing often assumes relatively deterministic behavior. Production AI evaluation requires a different approach: measure behavior continuously, detect statistically meaningful changes, and determine whether those changes originate from the AI system or from the evaluation process itself.
NiriZan is built around this principle.
It treats evaluation as an engineering process rather than a one-time benchmark.
NiriZan brings several capabilities together in one evaluation pipeline:
- Continuous AI evaluation
- LLM and RAG evaluation
- LLM-as-a-Judge evaluation
- Judge reliability analysis
- Judge drift attribution
- Fixed anchor sets and repeatable rescoring
- Statistical regression detection
- Statistical gating
- Multiple-comparison correction
- Bootstrap confidence intervals
- Trust-weighted health scoring
- CI/CD-integrated regression gates
The Wiki will gradually expand around several areas.
How NiriZan is structured and why its components are separated into distinct layers.
The ideas behind continuous evaluation, judge reliability, drift attribution, anchor sets, and related concepts.
The statistical methods used for regression detection, statistical gating, uncertainty estimation, and multiple-comparison correction.
Why NiriZan approaches AI evaluation as a continuous measurement and inference problem.
Practical information for using, developing, and contributing to NiriZan.
Answers to common questions about NiriZan's architecture, methodology, and implementation.
These sections are planned areas, not a claim that the corresponding Wiki pages are already complete.
The Wiki is not intended to replace the formal project documentation.
For installation instructions, API documentation, contracts, configuration, architecture specifications, and implementation details, see the official documentation:
For the source code, issues, development history, and contribution workflow:
For the published Python package:
The Wiki and formal documentation serve different purposes.
| Resource | Purpose |
|---|---|
| README | Project overview, discovery, and quick start |
| Wiki | Concepts, explanations, methodology, design rationale, and project knowledge |
| Documentation | Formal technical documentation, contracts, APIs, configuration, and implementation details |
| Source code | Authoritative implementation |
When the Wiki and implementation differ, the source code and formal project documentation take precedence for implementation behavior and contracts.
This Wiki is being built alongside NiriZan.
Some pages will eventually describe established functionality. Others may document research directions, experiments, architectural discussions, design decisions, or ideas that are still being evaluated.
The structure and content of the Wiki may therefore change as NiriZan develops.
This is the beginning, not the finished knowledge base.
π§ The Wiki is a work in progress.
The sections and links below represent the planned knowledge base. Some pages are already available, while others will be added as the Wiki develops.
Understand how NiriZan is structured and how its components interact.
Learn the concepts behind NiriZan before diving into implementation details.
- Continuous Evaluation
- AI Evaluation
- LLM-as-a-Judge
- Judge Reliability
- Judge Drift
- System Drift
- Judge Drift vs System Drift
- Anchor Sets
- Repeatable Rescoring
Understand how NiriZan uses statistical methods to make evaluation decisions.
- Statistical Gating
- Regression Detection
- Mann-Whitney U Testing
- Holm-Bonferroni Correction
- Bootstrap Confidence Intervals
- Effect Sizes
- Statistical Significance vs Practical Significance
Explore how NiriZan approaches evaluation as a continuous measurement problem.
- Why Continuous Evaluation?
- Why Statistical Evaluation?
- Evaluating Evaluators
- Judge Calibration
- Evaluation Reliability
Practical information for users and contributors.
Thank you for taking an interest in NiriZan.
Best regards,
Redwan Rahman