|
| 1 | +# TAGLINE |
| 2 | + |
| 3 | +Local GGUF LLM server with an OpenAI-compatible API |
| 4 | + |
| 5 | +# TLDR |
| 6 | + |
| 7 | +**Build** the Linux binary (downloads llama.cpp Vulkan shared libraries) |
| 8 | + |
| 9 | +```./build.sh``` |
| 10 | + |
| 11 | +**Start** the server (reads `.env` from the working directory or next to the binary) |
| 12 | + |
| 13 | +```janus``` |
| 14 | + |
| 15 | +**Start** without trying to open a browser |
| 16 | + |
| 17 | +```JANUS_NO_BROWSER=1 janus``` |
| 18 | + |
| 19 | +Run on **CPU** with an explicit GGUF file |
| 20 | + |
| 21 | +```INFERENCE_BACKEND=cpu JANUS_MODEL_PATH=[./models/model.Q4_K_M.gguf] janus``` |
| 22 | + |
| 23 | +**Listen** on a different address |
| 24 | + |
| 25 | +```JANUS_LISTEN_ADDR=[127.0.0.1:8991] janus``` |
| 26 | + |
| 27 | +**Check** that the process is up |
| 28 | + |
| 29 | +```curl http://127.0.0.1:8990/health``` |
| 30 | + |
| 31 | +Send an **OpenAI-compatible** chat request |
| 32 | + |
| 33 | +```curl http://127.0.0.1:8990/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"local","messages":[{"role":"user","content":"[Hello]"}]}'``` |
| 34 | + |
| 35 | +**Proxy** chat to a local Ollama instance |
| 36 | + |
| 37 | +```INFERENCE_BACKEND=ollama OLLAMA_MODEL=[mistral:7b] janus``` |
| 38 | + |
| 39 | +# SYNOPSIS |
| 40 | + |
| 41 | +**janus** |
| 42 | + |
| 43 | +# DESCRIPTION |
| 44 | + |
| 45 | +**janus** is a single Go binary from **Vibra-Ingenn** that loads a **GGUF** model on the local machine and serves it over HTTP. Inference uses **llama.cpp** through **Vulkan** (AMD, Intel, or NVIDIA) or a CPU fallback. The process also ships a bundled web UI (Assistant, Chat, Kernel, Config, Memory, Skills) and a **ReAct** tool loop that can read and write files, run shell commands, extract text from documents, and render PDF or Word output. |
| 46 | + |
| 47 | +There are no command-line flags. Startup reads a `.env` file (walking from the current directory toward the filesystem root, then next to the executable) and environment variables. The default listen address is **127.0.0.1:8990**. The OpenAI-style base URL is `http://127.0.0.1:8990/v1`. Useful paths include `/health`, `/v1/models`, `/v1/chat/completions` (streaming supported), `/v1/tools/list`, `/v1/tools/call`, `/kernel/run`, and `/upload` (50 MB multipart cap). |
| 48 | + |
| 49 | +On Linux, `./build.sh` fetches llama.cpp Vulkan `.so` files, compiles `dist/janus`, copies the libraries beside the binary, and sets `RPATH` to `$ORIGIN` when **patchelf** is available. Otherwise run with `LD_LIBRARY_PATH` pointing at the directory that contains `libllama.so`. The companion **modelget** binary in the same repository downloads a GGUF file from Hugging Face into `models/`. |
| 50 | + |
| 51 | +Hot-swap a model from the web UI Config tab or `POST /models/load` without restarting. Optional backends besides local GGUF are **Ollama** (`INFERENCE_BACKEND=ollama`) and cloud routers such as OpenRouter when the matching API key is set. |
| 52 | + |
| 53 | +# CONFIGURATION |
| 54 | + |
| 55 | +Copy `.env.example` to `.env` in the directory you launch **janus** from, or export the same names in the environment. Relative `JANUS_MODEL_PATH` values are resolved against the **current working directory**, not the source tree. |
| 56 | + |
| 57 | +**INFERENCE_BACKEND** |
| 58 | +> `vulkan` (default), `cpu`, `ollama`, or `openrouter`. |
| 59 | +
|
| 60 | +**JANUS_MODEL_PATH** |
| 61 | +> Path to the `.gguf` file. Required for local Vulkan/CPU inference. |
| 62 | +
|
| 63 | +**JANUS_LIB_PATH** |
| 64 | +> Path to `libllama.so` / `llama.dll`. Auto-detected next to the binary when unset. |
| 65 | +
|
| 66 | +**JANUS_GPU_LAYERS** |
| 67 | +> Layers to offload to the GPU. `-1` is all layers; `0` is CPU only. |
| 68 | +
|
| 69 | +**JANUS_VRAM_CEILING_MB** |
| 70 | +> Internal VRAM budget hint in MiB (default **9216**). |
| 71 | +
|
| 72 | +**JANUS_CTX_SIZE** |
| 73 | +> Context window in tokens. |
| 74 | +
|
| 75 | +**JANUS_MAX_TOKENS** |
| 76 | +> Maximum tokens per reply (default **4096**). |
| 77 | +
|
| 78 | +**JANUS_PROMPT_FORMAT** |
| 79 | +> Prompt template: `chatml`, `llama2`, or `alpaca`. |
| 80 | +
|
| 81 | +**JANUS_LISTEN_ADDR** |
| 82 | +> Bind address (default **127.0.0.1:8990**). A bare port is treated as `127.0.0.1:PORT`. |
| 83 | +
|
| 84 | +**JANUS_NO_BROWSER** |
| 85 | +> Set to any non-empty value to skip the automatic browser launch. |
| 86 | +
|
| 87 | +**JANUS_AUTH** |
| 88 | +> When `true`, require login on admin routes (Config, model load, and similar). |
| 89 | +
|
| 90 | +**JANUS_ADMIN_PASSWORD** |
| 91 | +> Admin password when auth is on. Generated on first run if unset. |
| 92 | +
|
| 93 | +**JANUS_SAFE_MODE** |
| 94 | +> When `true`, block shell commands from the `run_command` tool. |
| 95 | +
|
| 96 | +**JANUS_EXECUTION_MODE** |
| 97 | +> Kernel tool loop: `yolo` auto-runs tools; `safe` prompts before destructive tools. |
| 98 | +
|
| 99 | +**OLLAMA_BASE_URL** / **OLLAMA_MODEL** |
| 100 | +> Used when `INFERENCE_BACKEND=ollama` (default URL `http://127.0.0.1:11434`). |
| 101 | +
|
| 102 | +Logs append to `logs/janus.log`. Uploads land in `workspace/uploads/`; generated files in `workspace/outputs/`. Conversation facts persist in a local SQLite database. |
| 103 | + |
| 104 | +# CAVEATS |
| 105 | + |
| 106 | +This **janus** is the Vibra-Ingenn GGUF runner. Distro packages named **janus** are almost always the unrelated **Meetecho Janus WebRTC gateway**. Several other projects also ship a `janus` binary. |
| 107 | + |
| 108 | +The process takes no flags; a missing or wrong `.env` is the usual startup failure. Local Vulkan/CPU mode refuses to start without a readable `JANUS_MODEL_PATH`. Linux needs `libllama.so` beside the binary or on `LD_LIBRARY_PATH`. First model load commonly takes 10–60 seconds. The built-in browser opener uses Windows `cmd /c start` and does not open a browser on Linux. Built-in tools can run arbitrary shell commands unless **JANUS_SAFE_MODE** is set. OCR tools need **tesseract** on `PATH`. Default bind is loopback; do not expose the port without **JANUS_AUTH**. |
| 109 | + |
| 110 | +# HISTORY |
| 111 | + |
| 112 | +**Janus** is developed by **Vibra-Ingenn** as a MIT-licensed Go binary that wraps llama.cpp inference, an OpenAI-compatible router, and a local tool-using assistant. Windows is the primary target; Linux and macOS builds are supported from the same tree. |
| 113 | + |
| 114 | +# SEE ALSO |
| 115 | + |
| 116 | +[ollama](/man/ollama)(1), [llama-cli](/man/llama-cli)(1), [llama.cpp](/man/llama.cpp)(1), [llamafile](/man/llamafile)(1), [koboldcpp](/man/koboldcpp)(1), [vllm](/man/vllm)(1), [tesseract](/man/tesseract)(1) |
| 117 | + |
| 118 | +# RESOURCES |
| 119 | + |
| 120 | +```[Source code](https://github.com/Vibra-Ingenn/Janus)``` |
| 121 | + |
| 122 | +```[Documentation](https://github.com/Vibra-Ingenn/Janus/blob/main/docs/USER_MANUAL.md)``` |
| 123 | + |
| 124 | +<!-- verified: 2026-10-02 --> |
0 commit comments