Name and Version
bin/llama-cli --version
version: 7108 (4ae9216)
built with cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 for aarch64-linux-gn
Operating systems
Linux
GGML backends
CUDA, OpenCL
Hardware
rk3588 16GB memory
sudo cat /sys/kernel/debug/rknpu/version
RKNPU driver: v0.9.7
Models
https://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/tree/main
Problem description & steps to reproduce
bin/llama-server --verbose --mmproj ~/huggingface/Qwen3-VL-2B-Instruct-GGUF/mmproj-F16.gguf --host 0.0.0.0 -m ~/huggingface/Qwen3-VL-2B-Instruct-GGUF/Qwen3-VL-2B-Instruct-Q8_0.gguf
ask a question with image
First Bad Commit
No response
Relevant log output
main: server is listening on http://0.0.0.0:8080
main: starting the main loop...
que start_loop: processing new tasks
que start_loop: update slots
srv update_slots: all slots are idle
que start_loop: waiting for new tasks
add_text: <|im_start|>user
wtf
add_text: <|vision_start|>
image_tokens->nx = 11
image_tokens->ny = 3
batch_f32 size = 1
add_text: <|vision_end|>
add_text: <|im_end|>
<|im_start|>user
what is this
add_text: <|vision_start|>
image_tokens->nx = 5
image_tokens->ny = 4
batch_f32 size = 1
add_text: <|vision_end|>
add_text: <|im_end|>
<|im_start|>user
123
add_text: <|vision_start|>
image_tokens->nx = 5
image_tokens->ny = 5
batch_f32 size = 1
add_text: <|vision_end|>
add_text: <|im_end|>
<|im_start|>assistant
srv params_from_: Grammar:
srv params_from_: Grammar lazy: false
srv params_from_: Chat format: Content-only
srv add_waiting_: add task 0 to waiting list. current waiting = 0 (before add)
que post: new task, id = 0/1, front = 0
que start_loop: processing new tasks
que start_loop: processing task, id = 0
slot get_availabl: id 3 | task -1 | selected slot by LRU, t_last = -1
slot reset: id 3 | task -1 |
slot launch_slot_: id 3 | task -1 | launching slot : {"id":3,"n_ctx":4096,"speculative":false,"is_processing":false}
slot launch_slot_: id 3 | task -1 | sampler chain: logits -> logit-bias -> penalties -> dry -> top-n-sigma -> top-k -> typical -> top-p -> min-p -> xtc -> temp-ext -> dist
slot launch_slot_: id 3 | task 0 | processing task
que start_loop: update slots
srv update_slots: posting NEXT_RESPONSE
que post: new task, id = 1, front = 0
slot update_slots: id 3 | task 0 | new prompt, n_ctx_slot = 4096, n_keep = 0, task.n_tokens = 113
slot update_slots: id 3 | task 0 | n_tokens = 0, memory_seq_rm [0, end)
slot update_slots: id 3 | task 0 | prompt processing progress, n_tokens = 7, batch.n_tokens = 7, progress = 0.061947
srv update_slots: decoding batch, n_tokens = 7
clear_adapter_lora: call
set_embeddings: value = 0
Segmentation fault (core dumped)
Name and Version
bin/llama-cli --version
version: 7108 (4ae9216)
built with cc (Ubuntu 13.3.0-6ubuntu2~24.04.1) 13.3.0 for aarch64-linux-gn
Operating systems
Linux
GGML backends
CUDA, OpenCL
Hardware
rk3588 16GB memory
sudo cat /sys/kernel/debug/rknpu/version
RKNPU driver: v0.9.7
Models
https://huggingface.co/unsloth/Qwen3-VL-2B-Instruct-GGUF/tree/main
Problem description & steps to reproduce
bin/llama-server --verbose --mmproj ~/huggingface/Qwen3-VL-2B-Instruct-GGUF/mmproj-F16.gguf --host 0.0.0.0 -m ~/huggingface/Qwen3-VL-2B-Instruct-GGUF/Qwen3-VL-2B-Instruct-Q8_0.gguf
ask a question with image
First Bad Commit
No response
Relevant log output