Skip to content

Latest commit

 

History

70 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

VAGUE • Rust CLIP fastembed

vague is a local semantic search CLI engine. Searches match meaning instead of just text. A query like "that one legal file" returns a tax document even with zero word overlap. A query like "screenshot of an error message" returns a .png, since images are embedded into a vector space with CLIP instead of being tagged with just metadata. Text and image results are ranked together in a single list.

Public beta (v0.2.1): Stable build with all core features fully implemented. Maintained with occasional updates and feature additions.

Search Features

Text (.txt, .md, .pdf, .docx, .pptx, ...)

  1. File content is extracted and sent to a local text embedding model (nomic-embed-text) to produce a vector embedding.
  2. A search query is embedded the same way and compared against every indexed text file using cosine similarity.
vague index folder
vague search search "cat" --limit 1
[1.0000] \\?\D:\folder\animals.txt

use cases: finding specific documents and text files that have bad names, just by searching the context

Images (.png, .jpg, .jpeg, .webp)

  1. Images are embedded directly into vector space using CLIP, run locally via fastembed (Rust bindings over ONNX Runtime).
  2. Search queries are embedded through CLIP's own text encoder to land in the same space as the image vectors.
vague index folder
vague search "cat picture" --limit 1
[1.0000] \\?\D:\folder\cute_animals.png

use cases: finding a specific picture from just a description, looking through random screenshots with no text

OCR text detection in images (.png, .jpg, .jpeg, .webp)

  1. The optional --ocr flag extracts readable text from images during indexing.
  2. Extracted text is indexed alongside image vectors, so searches match words inside screenshots or diagrams. NOTE: Indexing with --ocr adds about ~0.4 seconds per image. Only use this if you need to search for text inside screenshots or diagrams.
vague index folder --ocr
vague search "tower-http" --limit 1
[1.3000] \\?\D:\folder\screenshot-logs.png

use cases: finding screenshots containing specific text, locating diagrams with labeled components

Installation

Windows

1. Download and extract

Download vague.zip from Releases and extract the folder to a stable location, like C:\tools\vague.

2. Add to PATH (optional, for terminal-wide access)

If you want to run vague from any terminal:

  • Open Settings and search "environment variables"
  • Click "Edit the system environment variables" -> "Environment Variables"
  • Under "User variables", select Path and click Edit (or create it if missing)
  • Click New and add C:\tools\vague
  • Click OK and restart your terminal

3. Verify

Run vague --help from any terminal to verify if the program is installed.

Linux/macOS

  1. Download vague.zip from Releases and extract it.
  2. Move the binary to your PATH and ensure the models folder is placed where vague expects it:
   sudo mv vague /usr/local/bin/
   sudo chmod +x /usr/local/bin/vague
   sudo mkdir -p /usr/local/share/vague
   sudo mv models/ /usr/local/share/vague/

(Note: If vague looks for the models folder relative to the executable's path, you can alternatively keep the binary and the models folder together in a dedicated directory like /opt/vague/ and symlink just the binary to /usr/local/bin/vague).

From Source (requires Rust + C++ dev tools 2022)

git clone https://github.com/ic0e/vague
cd vague
cargo install --path .

Updating

Download the latest vague.exe from Releases and replace the old one in your PATH folder (e.g., C:\tools).

Uninstall

Delete vague.exe from your PATH folder (e.g., C:\tools) and remove that folder from your PATH environment variable.

Usage

Setup & Basics

These are the first commands to run after installing vague.

vague --version   # prints the current version of vague
vague --help      # prints the current commands and their subcommands
vague setup       # downloads the text and image embedding models used for indexing

Running vague for the first time creates a settings file. Current settings:

  • limit - how many results search returns
  • depth - how deep indexing goes

To change a setting:

vague settings <setting> <value>

Example:

vague settings limit 10   # search now returns 10 results instead of the default 5
vague settings depth 1    # only indexes 1 directory deep

By default, the values are:

limit = 10 
depth = 9999 # a big number so it goes as deep as possible by default

Indexing

Files need to be indexed before they can be searched.

vague index /path/to/your/files

This generates embeddings and creates a searchable index. Models download automatically to ~/.vague_cache on first run (a few hundred MB, one time only).

Examples:

vague index .                       # indexes the cwd
cd pictures
vague index screenshots             # indexes a folder called screenshots
vague index screenshots/important   # indexes only that subfolder, not the rest of screenshots

Indexing can also run with OCR enabled, which extracts and saves text from images so it becomes searchable. This is slower than a normal index, so it's best used only on folders where the image text actually matters.

vague index <folder> --ocr

If you index without OCR first and want to add it later, you can run OCR separately afterward:

vague index <folder>
vague index <folder> --ocr   # only runs OCR on unindexed images, doesn't re-index existing files

Re-running index on a folder only picks up new files, it won't re-index files that are already in the index:

vague index .                 # indexes current folder
echo "text" > filename.txt    # adds a new file to the folder
vague index .                 # only indexes filename.txt, everything else isn't re-indexed

If existing files have changed and are no longer showing up correctly in search, the index needs to be overwritten instead. There are two ways to do this:

vague clear             # clears the entire index
vague index <folder>    # re-indexes from scratch

or in one step:

vague index . --overwrite   # clears and re-indexes in place, prompts with y/N

To change the depth of indexing just once without changing your settings:

vague index <folder> --depth <value>

Searching

vague search "find pictures of cats"
vague search "todo lists"

Results are ranked by relevance, both text and image matches in one output.

To see more or fewer results than the default without changing the setting permanently:

vague search "query" --limit 20   # returns 20 results for this search only

Requirements for development

  • Rust (2024 edition)
  • A C++ build toolchain — fastembed's ONNX Runtime bindings need to compile/link against C++ tooling.
    • Windows: install Visual Studio Build Tools with the "Desktop development with C++ 2022" workload selected. (make sure version is 2022)
    • macOS: Xcode Command Line Tools (xcode-select --install, etc.)
    • Linux: install build-essential (depends on distro, sudo apt install build-essential, etc.)

Development

To build and run from source:

git clone https://github.com/ic0e/vague.git
cd vague
cargo build --release
cargo run -- index <folder>
cargo run -- search "<query>"

OCR Support (for development)

OCR requires the detection and recognition models. To test OCR locally:

  1. Place text-detection.rten and text-recognition.rten in target/release/models/
  2. Run with --release only: cargo run --release -- index <folder> --ocr

OCR only works in release mode, debug builds makes it extremely slow due to ocrs (docs).

Project Layout For Devs

vague/
├── src/
│   ├── main.rs        # entry point, wires indexing and search together
│   ├── embedder.rs    # embeds text using `nomic-text-embed`
│   ├── clip.rs         # CLIP image + text-query embeddings via fastembed
│   ├── extract.rs     # reads text content from files
│   ├── indexer.rs      # walks a folder and builds the searchable index
│   └── store.rs        # normalized similarity scoring and merged ranked search

Roadmap & Future Features

  • Add showcase gifs in README
  • Video support (frame extraction + CLIP embedding per frame)
  • Further optimization of indexing and searching

Contributing

Please see CONTRIBUTING.md for guidelines on opening issues, testing features, and submitting pull requests.

License

This project is licensed under the GNU Affero General Public License v3.0 - see the LICENSE file for details.

About

A fast & local semantic search engine CLI in Rust. Multiple modules, searches text and images by meaning instead of exact words.

Topics

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages