Dataset for Training and Evaluating LLM-Based SOC Agents
-
Updated
Jun 4, 2026 - Python
Dataset for Training and Evaluating LLM-Based SOC Agents
Open-source dataset for evaluating Large Language Models (LLMs), developed as part of my graduation research project.
Grounded, fact-checked instruction-tuning dataset for cyber threat intelligence — 10k+ examples across 37 CTI categories, built from MITRE ATT&CK, CVE/KEV, CWE and live threat feeds. For LLM fine-tuning (LoRA/QLoRA).
Flipper Zero Sub-GHz RF dataset (280-1100 MHz, 9 countries) with 1500 Q&A pairs for LLM fine-tuning, fact-checked allocations, and a GPU-accelerated validation pipeline (Ollama Qwen 32B + DeBERTa NLI).
Open-source multi-server Discord channel scraper. Configure many servers in one TOML file, pick channels by name or ID, export to JSONL for offline search and LLM pipelines.
AI-powered Q&A system for U.S. affordable housing policy using RAG over 2,500+ HUD documents and 24 CFR
The Anti-Hallucination data layer for B2B Sourcing. Deep-verified global supply chain entities designed for RAG and LLM instruction tuning.
A comprehensive Python tool for extracting, processing, and analyzing RPG scenarios from the Era of the Imperial Republic (EOTIR) forums. Features automated web scraping, NLP-powered content analysis, character extraction, timeline generation, and LLM dataset preparation with an interactive HTML dashboard.
將維基文庫 (zh.wikisource.org) 下載的 EPUB/HTML 古籍一鍵轉換為乾淨 Markdown,自動識別並剝離導航欄、頁尾 noprint、姊妹計劃側欄、版權宣告與 MediaWiki 內部標記,保留純正文並注入結構化 YAML Front Matter(書名、卷號、來源)。支援生僻圖片字還原、雙行夾注轉全形括號、卷號異體字擴展。專為 LLM 訓練語料、RAG 向量資料庫與 Obsidian 個人知識庫建構而設計。僅依賴 Python 標準庫 + BeautifulSoup4,無需 C 編譯工具鏈,Termux / 樹莓派 / AWS Lambda 皆可零折騰部署。
A Curated RAG Dataset of 247 Articles on Chinese Muslim Food and Culture
This repository aims to provide a structured and easily accessible dataset of laws in Bangladesh. The data is primarily sourced from the Bangladesh Law (BDLAW) website.
Prepare the Kleister NDA dataset for LLM-based extraction. Validates labels against a Pydantic schema and delivers partitioned Parquet with co-located PDFs
1B-token JSONL training dataset mapping real neuroscience research (STDP, free energy, hippocampus, cortical columns, etc.) to disruptive software architecture paradigms. 100 paradigms × ~10K entries each.
Vet-reviewed Russian pet health Q&A: 3 709 owner question → veterinarian answer pairs (dogs, cats, exotics), species/category/source — non-synthetic (CC BY 4.0)
Gittxt is an AI-focused CLI and plugin tool for extracting, filtering, and packaging text from GitHub repos. Build LLM-compatible datasets, prep code for prompt engineering, and power AI workflows with structured .txt, .json, .md, or .zip outputs.
High-quality dataset of 201 authentic articles introducing Halal restaurants across China. RAG optimized.
Smart PDF-to-Dataset converter & chapter grouper for LLMs and NotebookLM. Converts large textbooks into clean JSON, CSV, and Excel with token optimization and auto-splitting.
Autonomous MCP server for M2M patent intelligence. Delivers structured JSON datasets (CPC A-H) enriched with biz_value_prop, tech stacks, and importance scoring. Supports instant autonomous data purchasing via ROSE cryptocurrency.
Features 232 articles covering Hui Muslim culture, travel, mosques, and halal food.
To associate your repository with the llm-dataset topic, visit your repo's landing page and select "manage topics."