AI-powered text simplification platform that transforms complex text into dyslexia-friendly formats using dual NLP pipelines — Normal and RAG-Enhanced.
DyslexiaLens leverages Natural Language Processing (NLP), Retrieval-Augmented Generation (RAG), and Google's Gemini AI to make text more accessible. The platform runs two parallel pipelines on any input text, allowing side-by-side comparison of simplification quality with comprehensive readability statistics.
- Normal NLP Pipeline — Lexical simplification, sentence splitting, grammar fixing, and Gemini-powered refinement
- RAG-Enhanced Pipeline — Same NLP core + ChromaDB-backed retrieval of expert dyslexia guidelines and vocabulary suggestions injected into the Gemini prompt
- 8 Readability Metrics — Flesch Reading Ease, Flesch-Kincaid Grade, Gunning Fog, SMOG Index, Coleman-Liau, Automated Readability Index, Dale-Chall Score, and Difficult Words Count
- Side-by-Side Comparison — View Original vs Normal vs RAG scores in a single table with delta indicators
- Issue Detection — Identifies long sentences, passive voice, and ambiguous structures in the original text
- Expert Rules Stream — Shows which dyslexia guidelines were retrieved from the vector database
- Lexical Suggestions Stream — Displays vocabulary replacements sourced from the RAG lexicon
- Issues Targeted — Lists the specific problems (passive voice, long sentences, complex vocabulary) that triggered rule retrieval
- Text Input — Paste or type text directly
- PDF Upload — Extracts text using PDFPlumber
- Image Upload — OCR via Tesseract with OpenCV preprocessing (PNG, JPG, JPEG)
- Dark-mode glassmorphism UI with smooth animations
- Animated loading states with pipeline status indicators
- Responsive design for desktop, tablet, and mobile
- Interactive tabs: Normal / RAG / Compare view modes
┌──────────────────────────────────────────────────────────┐
│ Frontend │
│ Next.js 15 + Vanilla CSS │
│ │
│ ┌──────────┐ ┌──────────┐ ┌────────────────────────┐ │
│ │ Input │ │ Pipeline │ │ Results Dashboard │ │
│ │ Section │→ │ Tabs │→ │ Stats · Table · Text │ │
│ └──────────┘ └──────────┘ └────────────────────────┘ │
└────────────────────────┬─────────────────────────────────┘
│ POST /simplif & POST /simplif-rag
▼
┌──────────────────────────────────────────────────────────┐
│ Backend (FastAPI) │
│ │
│ ┌────────────────────────────────────────────────────┐ │
│ │ Shared NLP Core │ │
│ │ clean → segment → readability → lexical simplify │ │
│ │ → sentence split → grammar fix → merge fragments │ │
│ └──────────────┬──────────────────┬──────────────────┘ │
│ │ │ │
│ ┌────────▼──────┐ ┌───────▼────────────────┐ │
│ │ Normal Path │ │ RAG Path │ │
│ │ │ │ │ │
│ │ Gemini with │ │ ChromaDB Retrieval: │ │
│ │ basic prompt │ │ ├─ Expert Rules DB │ │
│ │ │ │ └─ Lexicon DB │ │
│ │ │ │ ↓ │ │
│ │ │ │ Gemini with RAG prompt │ │
│ └───────┬───────┘ └────────┬────────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ JSON Response JSON Response │
│ (scores + text) (scores + text + rag_meta) │
└──────────────────────────────────────────────────────────┘
| Technology | Purpose |
|---|---|
| Next.js 15 | React framework with Turbopack |
| React 19 | Component-based UI |
| Vanilla CSS | Custom dark-mode design system |
| Axios | HTTP client for API calls |
| Technology | Purpose |
|---|---|
| FastAPI | Async Python web framework |
| NLTK | Tokenization and sentence splitting |
| spaCy | Linguistic analysis (passive voice detection) |
| TextStat | Readability score calculation |
| ChromaDB | Vector database for RAG retrieval |
| Google Gemini | AI-powered text refinement |
| Tesseract OCR | Image text extraction |
| PDFPlumber | PDF text extraction |
| ftfy | Text encoding cleanup |
DyslexiaLens/
├── backend/
│ └── main.py # FastAPI server with both pipelines
├── frontend/
│ ├── app/
│ │ ├── globals.css # Complete design system
│ │ ├── layout.tsx # Root layout with fonts & metadata
│ │ └── page.tsx # Main application page
│ ├── public/
│ ├── package.json
│ └── next.config.ts
├── notebook/
│ ├── nlp_pipeline.ipynb # Normal pipeline (research notebook)
│ ├── nlp_pipeline_rag.ipynb # RAG pipeline (research notebook)
│ └── extract.py # Notebook code extractor utility
├── .env # API keys (gemini_api_key, groq_api_key)
└── readme.md
- Python 3.10+
- Node.js 18+ and npm
- Tesseract OCR installed on your system
- Google Gemini API key
git clone https://github.com/prathoseraaj/DyslexiaLens.git
cd DyslexiaLens# Create and activate virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Install dependencies
pip install fastapi uvicorn python-multipart pydantic ftfy nltk spacy textstat \
google-generativeai python-dotenv pytesseract pillow pdfplumber opencv-python chromadb
# Download spaCy model
python -m spacy download en_core_web_smCreate a .env file in the project root:
gemini_api_key = YOUR_GEMINI_API_KEYcd frontend
npm installTerminal 1 — Backend:
source .venv/bin/activate
cd backend
python3 main.py
# or: uvicorn main:app --reload --host 0.0.0.0 --port 8000Terminal 2 — Frontend:
cd frontend
npm run dev- Frontend: http://localhost:3000
- Backend API: http://localhost:8000
- API Docs: http://localhost:8000/docs
| Method | Endpoint | Description |
|---|---|---|
GET |
/ |
Health check |
POST |
/simplif |
Normal NLP pipeline — returns simplified text + readability stats |
POST |
/simplif-rag |
RAG-enhanced pipeline — returns simplified text + stats + RAG metadata |
POST |
/upload |
Upload PDF/Image for text extraction |
{
"text": "Your complex paragraph here..."
}{
"simplified_text": "...",
"original_scores": {
"flesch_reading_ease": 28.5,
"flesch_kincaid_grade": 16.2,
"gunning_fog": 19.1,
"smog_index": 14.8,
"coleman_liau_index": 15.3,
"automated_readability_index": 17.9,
"dale_chall_readability_score": 10.2,
"difficult_words_count": 24,
"difficult_words_list": ["..."]
},
"final_scores": { "..." },
"improvement": 6.4,
"similarity": 0.72,
"assessment": {
"long_sentences": ["..."],
"passive_voice": ["..."],
"ambiguous_structures": ["..."]
},
"pipeline": "normal | rag",
"word_count_original": 120,
"word_count_simplified": 135,
"sentence_count_original": 6,
"sentence_count_simplified": 9,
"rag_metadata": {
"expert_rules_retrieved": ["..."],
"lexical_suggestions": ["..."],
"issues_detected": ["Passive Voice", "Long Sentences"],
"rules_count": 2,
"vocab_suggestions_count": 5
}
}- Text Cleaning — Fix encoding issues with
ftfy, normalize whitespace - Segmentation — Split into paragraphs, sentences, and tokens via NLTK
- Readability Analysis — Calculate 8 readability metrics using TextStat
- Issue Detection — Identify long sentences, passive voice (spaCy), and ambiguous structures
- Lexical Simplification — Replace 200+ complex words with simpler alternatives
- Sentence Splitting — Break sentences exceeding threshold at natural breakpoints
- Grammar Fixing — Clean up artifacts from simplification
- Fragment Merging — Rejoin sentence fragments starting with "And", "But", "Or"
- Gemini Refinement — Final AI pass to produce natural, meaning-preserving output
Adds two retrieval streams before the Gemini step:
- Expert Rules Retrieval — Queries a ChromaDB collection of 10 dyslexia accessibility rules based on detected issues (passive voice → "fix passive voice" query, etc.)
- Lexical Database Lookup — Queries the vocabulary vector store for the top 5 difficult words to find approved synonym mappings
- Augmented Gemini Prompt — Injects retrieved rules and vocabulary into the prompt with mandatory application instructions
| Metric | What It Measures | Target for Dyslexia |
|---|---|---|
| Flesch Reading Ease | Overall readability (0–100, higher = easier) | > 60 |
| Flesch-Kincaid Grade | US grade level needed to understand | < 8 |
| Gunning Fog | Years of education needed | < 10 |
| SMOG Index | Education years for 100% comprehension | < 10 |
| Coleman-Liau Index | Character-based grade level | < 8 |
| Automated Readability | Character & word-count based | < 8 |
| Dale-Chall Score | Familiar word ratio score | < 7 |
| Difficult Words | Count of unfamiliar words | Minimize |
This project is licensed under the Apache License 2.0.
Contributions are welcome! Please feel free to submit a Pull Request.
For questions or suggestions, please open an issue or contact the maintainer.
DyslexiaLens — Smart AI that reshapes text for dyslexic readers using dual NLP + RAG pipelines with full readability analytics.