VeriTruth is an AI-powered news verification web application. Users submit a news claim as text, a URL, or a screenshot/image, and VeriTruth returns a REAL / FAKE verdict produced by a fine-tuned BERT classifier, supported by transparent analysis signals and a human-readable AI explanation.
- Text verification — paste a news claim or article directly
- URL verification — article text is extracted from the link before analysis
- Image / screenshot verification — text is recovered from screenshots with Tesseract OCR
- BERT-based REAL / FAKE classification — the fine-tuned transformer is the main classifier
- Analysis Signals — source authority, similarity, recency, headline alignment, bias/tone and author signals computed from the submitted content
- Article-specific AI explanation — Groq generates a readable explanation of this article; it never changes the BERT verdict
- Feedback — users mark an analysis correct/incorrect with an optional reason
- Ask VeriTruth — a follow-up AI assistant (chat panel) that answers questions about the current analysis in context
- Authentication — email sign up / sign in, Google login, logout, and email-based password reset
- History & reports — saved analyses per user, plus an admin dashboard
| Layer | Technology |
|---|---|
| Backend | Python, Django 5.2 |
| Database | MySQL |
| ML model | PyTorch + Hugging Face Transformers (fine-tuned BERT) |
| OCR | Tesseract OCR (via pytesseract) |
| AI text | Groq API (explanations + assistant only — never classification) |
| Auth | Django auth + django-allauth (Google) + django-axes (lockout) |
| Frontend | HTML, CSS, JavaScript (vanilla) |
Text
Text → BERT → Result → Signals → AI Explanation
URL
URL → Article Extraction → BERT → Result → Signals → AI Explanation
Image
Image → Tesseract OCR → Extracted Text → BERT → Result → Signals → AI Explanation
- Feedback: every analysis can be rated correct/incorrect; ratings are stored for review in the admin dashboard. Feedback never retrains the model on its own.
- Ask VeriTruth: follow-up questions are sent — together with the current analysis context (verdict, confidence, signals, explanation) — to Groq, which answers in a scrollable chat panel without altering the original prediction.
BERT is always the decision-maker. Supporting signals contribute to the composite trust index, and Groq only explains and answers questions.
- Python 3.10+
- MySQL 8.x running locally
- Tesseract OCR installed
(Windows: add its folder to
PATHsotesseract --versionworks) - A Groq API key (free at console.groq.com)
- Optional: Google OAuth credentials for social login
git clone <your-repo-url>
cd fake_newsWindows:
python -m venv venv
venv\Scripts\activatemacOS / Linux:
python3 -m venv venv
source venv/bin/activatepip install -r requirements.txtDownload and install from the link above (skip if already installed).
CREATE DATABASE veritruth_db CHARACTER SET utf8mb4;copy .env.example .env # Windows
cp .env.example .env # macOS/LinuxFill in:
SECRET_KEY— generate one:python -c "from django.core.management.utils import get_random_secret_key; print(get_random_secret_key())"DB_NAME,DB_USER,DB_PASSWORD,DB_HOST,DB_PORTGROQ_API_KEYEMAIL_HOST_USER,EMAIL_HOST_PASSWORD(a Gmail app password for reset emails)
Google login: register the OAuth client in Google Cloud Console and add it in
Django admin → Social Applications (site: example.com for local dev), or set
GOOGLE_CLIENT_ID / GOOGLE_CLIENT_SECRET in .env.
The BERT weights are not stored in Git (the weights file is far larger than GitHub's 100 MB limit). See BERT Model Setup below for how to place or fetch the checkpoint.
Not committed on purpose: .env, model weights, training corpora under
detector/data/, media/ uploads, and local database dumps.
python manage.py migratepython manage.py createsuperuserpython manage.py runserver 127.0.0.1:8000model.safetensors is ~437 MB, which exceeds GitHub's 100 MB per-file limit, so
the weights are excluded from Git (*.safetensors is in .gitignore). No Git LFS
is required: the source lives on GitHub and the model is obtained separately from
the Hugging Face Hub.
Target layout — the app expects exactly these files:
ml/bert_fake_news/best_checkpoint/
├── config.json
├── model.safetensors
├── tokenizer.json
└── tokenizer_config.json
- Get a checkpoint. Either train locally
(
python ml/bert_training.py, which writesml/bert_fake_news/best_checkpoint/) or reuse the folder from an existing installation — an external drive or cloud storage works too. - Create a Hugging Face model repository at https://huggingface.co/new,
e.g.
YOUR_HUGGINGFACE_USERNAME/veritruth-bert-fake-news(public, so no token is needed to download it). - Upload the four checkpoint files above into the repository root — via the
web "Add files" UI or:
hf upload YOUR_HUGGINGFACE_USERNAME/veritruth-bert-fake-news ml/bert_fake_news/best_checkpoint
- Point the project at it. In
.env:BERT_MODEL_ID=YOUR_HUGGINGFACE_USERNAME/veritruth-bert-fake-news - On another machine, after cloning and installing requirements, fetch the
model into the checkpoint directory:
The script is idempotent: if the checkpoint is already complete it reports "Model already available" and downloads nothing.
python scripts/download_model.py
- Loading order.
ml/bert_service.pyuses the local checkpoint when all required files exist; otherwise it loadsBERT_MODEL_IDfrom the Hub, which downloads once and reuses the local cache afterwards. The model is never downloaded per request. - If neither source is available, analysis stops with a clear message — "VeriTruth BERT model is not available. Please complete the model setup described in the README." No other model is substituted and no placeholder verdict is produced.
- Private repository (optional). Set
HF_TOKENin the environment;huggingface_hubpicks it up automatically. Never commit a token.
Note: the
veritruth-bert-fake-newsrepository is a placeholder name in this README and in.env.example. The repository has not been created yet — replaceYOUR_HUGGINGFACE_USERNAMEonce you publish it.
fake_news/
├── manage.py
├── fake_news/ # Django project (settings, urls, wsgi)
├── detector/ # Main app: views, models, auth, OCR, Groq integration
│ └── migrations/
├── ml/ # BERT inference service + model checkpoints
├── scripts/ # Setup helpers (download_model.py)
├── templates/ # HTML templates
├── static/ # CSS / JS / images
├── media/ # User uploads (runtime, not committed)
└── requirements.txt
- All credentials (database, Groq, SMTP, OAuth,
SECRET_KEY) are loaded from.env— never hard-coded and never committed. - Login attempts are rate-limited by django-axes (5 failures → temporary lockout).
- Analysis routes require authentication; logged-out visitors see a compact sign-in prompt before submitting.
- Password-reset links only reach accounts whose stored
emailis real and deliverable. Django deliberately shows the same "reset link sent" success page for an unknown or mistyped address (so the form cannot be used to enumerate accounts), so that page is not proof of delivery — confirm against the mailbox itself.
VeriTruth is a research/prototype tool. Its verdicts are probabilistic model predictions, not ground truth. Always verify important claims through multiple reliable sources.