Technical aspects of running local LLMs on FreeBSD

There is no FreeBSD-specific CUDA code, obviously.

we can agree on that
makes a change

As you said ports have to be patched

And there are some ports that havent been patched like Handbrake or Blender,
because you only have so much time, and as you said about Handbrake

FYI, there is no NVENC with Handbrake, because I didn't bother to submit the corresponding patches to the port.

That's correct isnt it.

So the project is addressing 2 issues

1) Ports that havent been patched to work with Cuda,
and installing the Linux version with Podman to get Cuda working.

2) Linux applications that require Cuda but cant be ported to Freebsd like Davinci Resolve because of licensing issues.

So if a port hasnt been patched the user has an alternative,
and can install the Linux version of the application with Podman and get it working with Cuda.

Rather than asking you or other devs to patch lots of different ports,
which you might not have the time for.

Surely thats a good thing from your point of view,
as it means less work patching.

Having pre configured applications that are easy to deploy,
and making things easier for user and giving them more choice is a good thing.

The end goal is getting an application working with Cuda for users,
we just have different ways of going about it.
 
ThisIsAGoodThread.png
 
To run a application/port that has been patched you have to prefix the command with nv-sglrun

for example to use ffmpeg with nvenc

Code:
nv-sglrun ffmpeg

script example

Code:
#!/bin/sh

#===============================================================================
# convert video to h265/aac
#===============================================================================


#===============================================================================
# script usage
#===============================================================================

usage()
{
# if argument passed to function echo it
[ -z "${1}" ] || echo "! ${1}"
# display help
echo "\
# convert video to h265/aac

$(basename "$0") -i input.mov -o output.mp4
-i infile.mov
-o outfile.mov :optional agument # if option not provided defaults to input-name.mp4"
exit 2
}


#===============================================================================
# error messages
#===============================================================================

NOTFILE_ERR='not a file'
INVALID_OPT_ERR='Invalid option:'
REQ_ARG_ERR='requires an argument'
WRONG_ARGS_ERR='wrong number of arguments passed to script'


#===============================================================================
# check number of aruments passed to script
#===============================================================================

[ $# -gt 0 ] || usage "${WRONG_ARGS_ERR}"


#===============================================================================
# getopts check options passed to script
#===============================================================================

while getopts ':i:o:h' opt
do
  case ${opt} in
     i) input="${OPTARG}"
    [ -f "${input}" ] || usage "${input} ${NOTFILE_ERR}";;
     o) output="${OPTARG}";;
     h) usage;;
     \?) usage "${INVALID_OPT_ERR} ${OPTARG}" 1>&2;;
     :) usage "${INVALID_OPT_ERR} ${OPTARG} ${REQ_ARG_ERR}" 1>&2;;
  esac
done
shift $((OPTIND-1))


#===============================================================================
# variables
#===============================================================================

input_nopath="${input##*/}"
input_name="${input_nopath%.*}"

# defaults for variables if not defined
output_default="${input_name}.mp4"


#===============================================================================
# functions
#===============================================================================

# h265 function
h265 () {
    nv-sglrun \
    ffmpeg \
    -hide_banner \
    -stats -v panic \
    -i "${input}" \
    -c:v hevc_nvenc \
    -pix_fmt p010le \
    -preset slow \
    -tier high \
    -rc vbr \
    -cq 22 \
    -b:v 0 \
    -maxrate 50M \
    -c:a aac \
    -b:a 320k \
    -ar 48000 \
    "${output:=${output_default}}"
}

# run the h265 function
h265 "${input}"

for gui applications you need to modify the desktop entry

for example to get obs studio working with nvenc

Code:
[i] Yes Master ? ls -l ~/.local/share/applications/com.obsproject.Studio.desktop
-rw-r--r--  1 djwilcox djwilcox 395  6 Jul 15:04 /home/djwilcox/.local/share/applications/com.obsproject.Studio.desktop

com.obsproject.Studio.desktop

Code:
[Desktop Entry]
Version=1.0
Name=OBS
GenericName=Streaming/Recording Software
Comment=Free and Open Source Streaming/Recording Software
Exec=sh -c 'LD_LIBMAP="`nv-sglrun printenv LD_LIBMAP | grep -v libGL`" obs --websocket_ipv4_only'
Icon=com.obsproject.Studio
Terminal=false
Type=Application
Categories=AudioVideo;Recorder;
StartupNotify=true
StartupWMClass=obs

notice the exec line

Code:
Exec=sh -c 'LD_LIBMAP="`nv-sglrun printenv LD_LIBMAP | grep -v libGL`" obs --websocket_ipv4_only'

as opposed to just running obs

Code:
Exec=obs --websocket_ipv4_only

However as i said if you try and install something like whisperx
which doesnt have a native Freebsd package with pip or conda with the Linuxulator or a Jail

Then it will fail to install because it will detect its running on Freebsd
and try and download Freebsd python wheels that dont exist

there are Freebsd python packages for torch

Code:
[i] Yes Master ? pkg search torch
py312-facenet-pytorch-2.5.3_4  Pretrained PyTorch face detection and recognition models
py312-lion-pytorch-0.2.4       PyTorch: Lion optimizer
py312-pytorch-2.12.1           PyTorch: Tensors and dynamic neural networks in Python
py312-pytorch-lightning-2.6.5  Lightweight PyTorch wrapper for ML researchers
py312-pytorchvideo-0.1.5_4     Video understanding deep learning library
py312-torch-geometric-2.8.0    Graph neural network library for PyTorch
py312-torchao-0.17.0           PyTorch: Package for applying ao techniques to GPU models
py312-torchaudio-2.11.0        PyTorch-based audio signal processing and machine learning library
py312-torchcodec-0.13.0        PyTorch media decoding and encoding
py312-torchdata-0.11.0         PyTorch: Composable data loading modules for PyTorch
py312-torchmetrics-1.9.0       PyTorch native metrics
py312-torchsde-0.2.6_2         SDE solvers and stochastic adjoint sensitivity analysis in PyTorch
py312-torchsummary-1.5.1_2     PyTorch: Model summary in PyTorch
py312-torchvision-0.27.0       PyTorch: Datasets, transforms and models specific to computer vision
pytorch-2.12.1                 Tensors and dynamic neural networks in Python (C++ library)
[

But they may not be the correct version for python application you are trying to install
and the application may require additional python libraries that dont have a Freebsd package

So the advantage of using Podman is you can set the container platform and os to Linux
which presents a Linux environment to applications like python so they download the python linux wheels

So for native Freebsd packages that have been patched to support nvenc for example
you can prefix the command with nv-sglrun or modify the desktop entry
 
BSD Jedi has some Ollama tutorial and a new one with a plug for some budget option cloud models.

I run 16GB vRAM nvidia on Debian Linux Ollama - rarely used, mainly as brain for Home Assistant....for fun.
I have good experience with Mixture of Experts models - gemma4. They really seem to use just parts of the model they actually need for given task.

Alas at work I have some good Ryzen AMD CPU, no vRAM to speak of but 64GB of DDR5 RAM (bought a year ago, yey) and it is surprisingly capable of running much bigger models than my home machine. Slow but who cares when crunching config files and not having conversations.

I also bought Google Coral TPUs on cheap but there is no chance I can make them work on naked FreeBSD - python libraries expect Linux. I was thinking of making a robot car with visual recognition on FreeBSD but all the python libraries are just made for Linux Rapsberry "drivers" too.

For more serious and private work 128GB is a must. But as OP pointed out if privacy is not a must, cloud models are better investment. Or rather 4,000 USD machine which is way dumber is just really bad investment right now.
 
Edit: 2026-09-26 - updated to make llama run by user

Some quick and crude steps to get llama-cpp (version: 0.4.0-dev (build 10931, commit 3057bb66c) running with CUDA in the Linuxlator on FreeBSD 15.1.


llama-cpp with CUDA on the Linuxlator

# ----------------------------------------------------------------
# The FreeBSD side
# ----------------------------------------------------------------
# The necessary filesystems
# Note: Last line only necessary if you have models already
# on the BSD side. Change it to your path.

/etc/fstab
proc /proc procfs rw 0 0
linprocfs /compat/linux/proc linprocfs rw 0 0
linsysfs /compat/linux/sys linsysfs rw 0 0
tmpfs /compat/linux/dev/shm tmpfs rw,late,mode=1777 0 0
/home/<insertyourusernamehere>/AI-models /compat/linux/home/<insertyourusernamehere>/AI-models nullfs rw,late 0 0

# /etc/rc.conf
sudo sysrc linux_enable="YES"
sudo sysrc kld_list+="nvidia-modeset nvidia-drm linux64"

# Note: these are the drivers tested
nvidia-driver-595.99.02 NVIDIA graphics driver userland
nvidia-drm-612-kmod-595.99.02.1501000 NVIDIA DRM Kernel Module
nvidia-drm-kmod-595.99.02 NVIDIA DRM kernel module
nvidia-kmod-595.99.02.1501000 NVIDIA graphics driver kernel module

# Setting up the Linuxlator
sudo mkdir -p /compat/linux
sudo debootstrap --arch=amd64 jammy /compat/linux http://archive.ubuntu.com/ubuntu
sudo mount /compat/linux/proc
sudo mount /compat/linux/sys
sudo mount /compat/linux/dev/shm
sudo service linux start

# Below is version 595.99.02
cd /usr/ports/x11/linux-nvidia-libs
sudo make install
sudo mkdir -p /compat/linux/home/<insertyourusernamehere>/AI-models
sudo mount -t nullfs /home/<insertyourusernamehere>/AI-models /compat/linux/home/<insertyourusernamehere>/AI-models

# ----------------------------------------------------------------
# The Linux side
# ----------------------------------------------------------------
# Initializing Ubuntu
sudo chroot /compat/linux /bin/bash
apt update
apt-get install -y software-properties-common && add-apt-repository universe && apt update
apt install -y python3-pip python3-dev wget curl build-essential
strings /usr/lib/x86_64-linux-gnu/libstdc++.so.6 | grep GLIBCXX_3.4.30
apt install -y ninja-build cmake
mkdir -p /opt/llama/bin

# The FreeBSD uvm modification
mkdir /tmp/freebsd-cuda
cd /tmp/freebsd-cuda
git clone https://github.com/NapoleonWils0n/freebsd-cuda.git .
cd /opt/llama/bin
cp /tmp/freebsd-cuda/base-build/uvm_ioctl_override/uvm_ioctl_override.c .
gcc -shared -fPIC -o /opt/llama/bin/dummy-uvm.so /opt/llama/bin/uvm_ioctl_override.c

# Installing cuda stuff
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-ubuntu2204.pin
mv cuda-ubuntu2204.pin /etc/apt/preferences.d/cuda-repository-pin-600
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.1-1_all.deb
dpkg -i cuda-keyring_1.1-1_all.deb
apt update
# These three below might come with cuda-toolkit as well
apt install -y cuda-cudart-12-8
apt install -y libcublas-12-8
apt install -y libnccl2

# Builing llama-cpp
cd /opt
git clone https://github.com/ggerganov/llama.cpp.git llama-build
cd llama-build
apt install -y cuda-compiler-12-8 cuda-toolkit-12-8
export PATH="/usr/local/cuda-12.8/bin:$PATH"
echo 'export PATH="/usr/local/cuda-12.8/bin:$PATH"' >> /root/.bashrc
export CUDAToolkit_ROOT="/usr/local/cuda-12.8"
echo 'export CUDAToolkit_ROOT="/usr/local/cuda-12.8"' >> /root/.bashrc
# Fix a 32/64 bit library issue
rm -f /usr/lib/libcuda.so
ln -s /usr/local/cuda-12.8/targets/x86_64-linux/lib/stubs/libcuda.so /usr/lib/libcuda.so
cd /opt/llama-build
rm -rf build
cmake -B build -G Ninja -DGGML_CUDA=ON
cmake --build build --config Release --target llama-server


# Syncing user FreeBSD side/Linux side
grep "^insertyourusernamehere:" /etc/passwd | sudo tee -a /compat/linux/etc/passwd

# Running llama-cpp with cuda on FreeBSD
PATH="/compat/linux/usr/local/cuda-12.8/bin:$PATH" \
CUDAToolkit_ROOT="/compat/linux/usr/local/cuda-12.8" \
LD_LIBRARY_PATH
LD_PRELOAD="$DUMMY_UVM_HOOK" \

env LD_LIBRARY_PATH="/compat/linux/usr/local/cuda-12.8/lib64:/usr/lib64:/usr/lib/x86_64-linux-gnu" \
LD_PRELOAD="/compat/linux/opt/llama/bin/dummy-uvm.so" \
/compat/linux/opt/llama-build/build/bin/llama-server \
-m /home/<insertyourusernamehere>/AI-models/qwen2.5-coder-14b-instruct-q4_k_m.gguf \
--host 0.0.0.0 \
--port 8080 \
-ngl 99


Note: Running nvidia-smi on the Linux side will not show CUDA - this is most likely because nvidia-smi not is built with the uvm modification.

Here's are some snippets of the output when llama-cpp start up:


0.00.014.956 I cmn common_param: common_params_print_info: build 10931 (3057bb66c) with GNU 11.2.0 for Linux x86_64
0.00.014.961 I cmn common_param: common_params_print_info: verbosity = 6 (adjust with the `-lv N` CLI arg)
0.00.014.961 I cmn common_param: device_info:
0.00.152.558 I cmn common_param: - CUDA0 : NVIDIA GeForce RTX 3060 (12157 MiB, 11603 MiB free)
0.00.152.568 I cmn common_param: - CPU : Intel(R) Core(TM) i5-2500 CPU @ 3.30GHz (16297 MiB, 16297 MiB free)
0.00.152.688 I cmn common_param: system_info: n_threads = 4 (n_threads_batch = 4) / 4 | CUDA : ARCHS = 500,610,700,750,800,860,890,900,1200 | USE_GRAPHS = 1 | FA_QUANTS = q4_0-q4_0,q8_0-q8_0,f16-f16,bf16-bf16 | CPU : SSE3 = 1 | SSSE3 =
1 | AVX = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |
0.00.152.695 I srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
0.00.152.825 I srv init: using 8 threads for HTTP server

....

0.00.264.148 I print_info: file size = 8.37 GiB (4.87 BPW)
0.00.264.669 I llama_prepare_model_devices: using device CUDA0 (NVIDIA GeForce RTX 3060) (0000:01:00.0) - 11603 MiB free
0.00.357.956 D init_tokenizer: initializing tokenizer for type 2
0.00.392.019 I load: 0 unused tokens

....

0.00.515.627 I print_info: max token length = 256
0.00.515.872 I load_tensors: loading model tensors, this can take a while... (load_mode = none)
0.00.516.326 D load_tensors: layer 0 assigned to device CUDA0, is_swa = 0
0.00.516.327 D load_tensors: layer 1 assigned to device CUDA0, is_swa = 0
0.00.516.328 D load_tensors: layer 2 assigned to device CUDA0, is_swa = 0
0.00.516.328 D load_tensors: layer 3 assigned to device CUDA0, is_swa = 0


/grandpa
 
Last edited:
Some more technical aspects.

Running llama-cpp with CUDA on the Linuxlator and aider on the FreeBSD side with the same prompt (see #10) . Still on crap hardware.

The random models and how they did.

#SizeNameResultCUDAVULKAN
1)4.36Gqwen2.5-coder-7b-instruct-q4_k_m.ggufsuccess for load average but fail for temperature~ 62 t/s~ 54 t/s
2)
3)6.87Ggemma4-coding-Q4_K_M.ggufsuccess~ 33 t/s~ 30 t/s
4)
5)8.01GCodestral-22B-v0.1-IQ3_XXS.ggufsuccess~ 24 t/s~ 18 t/s
6)8.37Gqwen2.5-coder-14b-instruct-q4_k_m.ggufsuccess for temperature fail for load averages~ 32 t/s~ 30 t/s
7)
8)8.65GDevstral-Small-2-24B-Instruct-2512-UD-Q2_K_XL.ggufsuccess for temperature fails for load average~ 21 t/s~19 t/s
9)8.74GNorth-Mini-Code-1.0-UD-IQ1_M.ggufsuccess~ 72 t/s~ 54 t/s
a)
b)9.11Ggemma4-coding-Q6_K.ggufsuccess~ 27 t/s~25 t/s
c)9.11GQwen3.5-9B-Q8_0.gguffails stuck in a loop~ 34 t/s~25 t/s
d)14.84GQwen2.5-Coder-32B-Instruct-Q3_K_M.ggufsuccess for temperature fails slightly for load averages (2/3)~ 1.22 t/s~ 1 t/s

Notes:
-
Results may vary. Maybe it's the tweaking of the parameters. Especially suspect is the context window size.
- The feeling from #10 that Vulkan not always is giving memory back is totally gone. Switching between models does not cause out of device memory.
- In general it feels more smooth.

/grandpa
 
Edit 2026-10-04: The original post was removed and replaced with this.

For llama-server running on Linuxlator with CUDA here are some scripts to calculate maximum context window and then run aider or claude.


These scripts attempts to:
  • check computer RAM and NVRAM
  • read a list of models and their size from a specified directory
  • calculate maximum context window options for llama-cpp to fit inside the hardware constraints
  • start aider with parameters coherent with llama-cpp or
  • start claude with parameters coherent with llama-cpp
  • make it possible to run models that are larger than VRAM by offloading to RAM
First a disclaimer.
- To get CUDA working with Linuxlator look at #30
- This post became so large I had to split it into #32, #33, #34 and #35. Sorry about that. It's 11 scripts altogether.
- The scripts below are just a snapshot of the moment in time when this is being written. The evolution of llama-cpp and aider is going fast so expect change.
- The Scripts uses experimental flags like "tools all" for llama-cpp - be warned

This is the software from this snapshot in time:

  • FreeBSD 15.1-RELEASE-p3 releng/15.1-n283611-88e7371d9dc2 GENERIC amd64 including python3.12
  • aider 0.86.3.dev53+g5dc9490bb.d20260902 (pip installed in a venv and a modified requirements.txt)
  • linuxlator and CUDA installed as described in #30
  • llama-cpp version: 0.4.0-dev (build 10931, commit 3057bb66c) installed as in #30
  • some AI-models from huggingface
  • claude 2.1.281 (Claude Code) installed as described in #34

How do I run these scripts?
  1. Put the scripts in your shell path and make sure they are executable
  2. run-ai.sh
Note: A file name .git-analyzer.conf will be created and saved in user $HOME directory.

The scripts:

sh:
#!/bin/sh

# ==============================================================================
# 🌐 AI max context calculator
# ==============================================================================
CONFIG_FILE="$HOME/.git-analyzer.conf"

unset CONFIG_SCRIPTS_PATH CONFIG_MODELS_PATH CONFIG_WORKSPACE_PATH
unset CONFIG_AIDER_PATH CONFIG_CLAUDE_PATH CONFIG_SERVER_HOST CONFIG_SERVER_PORT
unset CONFIG_LLAMA_SERVER_BIN CONFIG_DUMMY_UVM_HOOK CONFIG_VRAM_OVERHEAD CONFIG_RAM_OVERHEAD
unset CONFIG_CONTEXT_MULTIPLIER CONFIG_DEBUG

# If the configuration file is missing at first launch, run the parameters setup directly
if [ ! -f "$CONFIG_FILE" ]; then
    echo "⚠️  No configuration profile detected. Launching global parameters setup..."
    if [ -f "$HOME/scripts/config-params.sh" ]; then
        "$HOME/scripts/config-params.sh"
    elif [ -f "./config-params.sh" ]; then
        "./config-params.sh"
    else
        echo "❌ Critical Error: config-params.sh not found in ~/scripts or current directory." >&2
        exit 1
    fi
    if [ $? -ne 0 ]; then exit 1; fi
fi

. "$CONFIG_FILE"

SCRIPTS_DIR="$CONFIG_SCRIPTS_PATH"
MODELS_DIR="$CONFIG_MODELS_PATH"
WORKSPACE_DIR="$CONFIG_WORKSPACE_PATH"
AIDER_BIN_PATH="$CONFIG_AIDER_PATH"
CLAUDE_BIN_PATH="$CONFIG_CLAUDE_PATH"
SERVER_HOST_VAL="$CONFIG_SERVER_HOST"
SERVER_PORT_VAL="$CONFIG_SERVER_PORT"
LLAMA_SERVER_BIN="$CONFIG_LLAMA_SERVER_BIN"
DUMMY_UVM_HOOK="$CONFIG_DUMMY_UVM_HOOK"
VRAM_OVERHEAD_GB="${CONFIG_VRAM_OVERHEAD:-1.5}"
RAM_OVERHEAD_GB="${CONFIG_RAM_OVERHEAD:-2.0}"
CONTEXT_MULTIPLIER="${CONFIG_CONTEXT_MULTIPLIER:-1.5}"
DEBUG_MODE="${CONFIG_DEBUG:-false}"

get_idx_label() {
    val="$1"
    if [ "$val" -le 9 ]; then echo "${val})"; else
        case "$val" in
            10) echo "a)" ;; 11) echo "b)" ;; 12) echo "c)" ;;
            13) echo "d)" ;; 14) echo "e)" ;; 15) echo "f)" ;;
            *)  echo "?)" ;;
        esac
    fi
}

while true; do
echo "──────────────────────────────────────────────────────"
echo "🌐 AI max context calculator"
echo "──────────────────────────────────────────────────────"
echo "💡 Path to scripts: ${SCRIPTS_DIR}"
echo "💡 Path to AI Models: ${MODELS_DIR}"
echo "💡 Path to workspace: ${WORKSPACE_DIR}"
echo "💡 Path to Aider binary: ${AIDER_BIN_PATH}"
echo "💡 Path to Claude binary: ${CLAUDE_BIN_PATH}"
echo "💡 Llama host ip: ${SERVER_HOST_VAL}"
echo "💡 Llama host port: ${SERVER_PORT_VAL}"
echo "💡 Path to llama binary: ${LLAMA_SERVER_BIN}"
echo "💡 Path to dummy-uvm-hook: ${DUMMY_UVM_HOOK}"
echo "💡 VRAM overhead: ${VRAM_OVERHEAD_GB}"
echo "💡 Ram overhead: ${RAM_OVERHEAD_GB}"
echo "💡 Trained Context Multiplier (X): ${CONTEXT_MULTIPLIER}x"
echo "💡 Global Verbose Debug Mode: ${DEBUG_MODE}"
echo "──────────────────────────────────────────────────────"
echo "Select execution path:"
echo "1) List and select model"
echo "2) Change configuration paths / network settings"
echo "0) Exit"
echo "──────────────────────────────────────────────────────"
printf "Your choice: "
read -r menu_choice

case "$menu_choice" in
    1)
        echo "🔍 Scanning models path..."
        echo "────────────────────────────────────────────────────────────────────────────────────────"
        printf "%-5s | %-55s | %-10s\n" "IDX" "MODEL NAME" "SIZE"
        echo "────────────────────────────────────────────────────────────────────────────────────────"
     
        RAW_MATRIX_FILE="/tmp/matrix_raw.tmp"
        rm -f "$RAW_MATRIX_FILE"
     
        for model_file in "$MODELS_DIR"/*.gguf; do
            if [ -f "$model_file" ]; then
                M_NAME=$(basename "$model_file")
                BYTES=$(stat -f %z "$model_file" 2>/dev/null || stat -c %s "$model_file")
                M_SIZE=$(awk -v b="$BYTES" 'BEGIN {printf "%.2f GB", b / 1024 / 1024 / 1024}')
                ROW_CONTENT=$(printf "%-55.55s | %-10s" "$M_NAME" "$M_SIZE")
                echo "$BYTES    $ROW_CONTENT    $model_file" >> "$RAW_MATRIX_FILE"
            fi
        done
     
        SORTED_LIST_FILE="/tmp/selected_models.list"
        rm -f "$SORTED_LIST_FILE"
     
        counter=1
        if [ -f "$RAW_MATRIX_FILE" ]; then
            sorted_data=$(sort -n "$RAW_MATRIX_FILE")
            rm -f "$RAW_MATRIX_FILE"
            echo "$sorted_data" | while IFS='    ' read -r bytes row_content full_path; do
                if [ -n "$row_content" ]; then
                    idx_lbl=$(get_idx_label "$counter")
                    printf "%-5s | %s\n" "$idx_lbl" "$row_content"
                    echo "$full_path" >> "$SORTED_LIST_FILE"
                    counter=$((counter + 1))
                fi
            done
        fi
        echo "────────────────────────────────────────────────────────────────────────────────────────"
        printf "Select model or 0 to return: "
        read -r model_choice
     
        if [ "$model_choice" != "0" ] && [ -n "$model_choice" ]; then
            case "$model_choice" in
                a) target_row=10 ;; b) target_row=11 ;; c) target_row=12 ;;
                d) target_row=13 ;; e) target_row=14 ;; f) target_row=15 ;;
                *) target_row="$model_choice" ;;
            esac
         
            SELECTED_MODEL_PATH=$(sed -n "${target_row}p" "$SORTED_LIST_FILE" 2>/dev/null)
         
            if [ -n "$SELECTED_MODEL_PATH" ] && [ -f "$SELECTED_MODEL_PATH" ]; then
                M_NAME_FINAL=$(basename "$SELECTED_MODEL_PATH")
                FILE_SIZE_BYTES=$(stat -f %z "$SELECTED_MODEL_PATH" 2>/dev/null || stat -c %s "$SELECTED_MODEL_PATH")
                MODEL_SIZE_GB=$(awk -v b="$FILE_SIZE_BYTES" 'BEGIN {printf "%.2f", b / 1024 / 1024 / 1024}')

                # 🎯 Execute context computation scripts
                VRAM_DATA=$("$SCRIPTS_DIR/calc-ctx_vram.sh" "$SELECTED_MODEL_PATH")
                RAM_DATA=$("$SCRIPTS_DIR/calc-ctx_ram.sh" "$SELECTED_MODEL_PATH")

                # 🎯 Secure field parsing and single-line isolation using cut and head
                CTX_V_F16=$(echo "$VRAM_DATA" | head -n1 | cut -d'|' -f1 | tr -d '\n\r ')
                CTX_V_Q8=$(echo "$VRAM_DATA" | head -n1 | cut -d'|' -f2 | tr -d '\n\r ')
                CTX_V_Q4=$(echo "$VRAM_DATA" | head -n1 | cut -d'|' -f3 | tr -d '\n\r ')

                NGL_SPILL=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f1 | tr -d '\n\r ')
                RAM_USED_BY_MODEL=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f2 | tr -d '\n\r ')
                CTX_R_F16=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f3 | tr -d '\n\r ')
                CTX_R_Q8=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f4 | tr -d '\n\r ')
                CTX_R_Q4=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f5 | tr -d '\n\r ')
                TOTAL_LAYERS=$(echo "$RAM_DATA" | head -n1 | cut -d'|' -f6 | tr -d '\n\r ')

                # 🐛 CONDITIONAL DEBUG PRINTOUTS
                if [ "$DEBUG_MODE" = "true" ]; then
                    echo "🐛 [DEBUG run-ai.sh] Model: $M_NAME_FINAL (Bytes: $FILE_SIZE_BYTES)" >&2
                    echo "🐛 [DEBUG run-ai.sh] VRAM -> F16: $CTX_V_F16 | Q8: $CTX_V_Q8 | Q4: $CTX_V_Q4" >&2
                    echo "🐛 [DEBUG run-ai.sh] RAM  -> Layers: $TOTAL_LAYERS | NGL: $NGL_SPILL | Spill: $RAM_USED_BY_MODEL GB" >&2
                fi

                # Mark if model goes OOM (Out Of memory)
                [ -z "$CTX_R_F16" ] || [ "$CTX_R_F16" -eq 0 ] 2>/dev/null && DISP_R_F16="N/A (OOM Risk)" || DISP_R_F16="${CTX_R_F16} tokens"
                [ -z "$CTX_R_Q8" ]  || [ "$CTX_R_Q8" -eq 0 ]  2>/dev/null && DISP_R_Q8="N/A (OOM Risk)"  || DISP_R_Q8="${CTX_R_Q8} tokens"
                [ -z "$CTX_R_Q4" ]  || [ "$CTX_R_Q4" -eq 0 ]  2>/dev/null && DISP_R_Q4="N/A (OOM Risk)"  || DISP_R_Q4="${CTX_R_Q4} tokens"
                echo ""
                echo "========================================================================================"
                echo "📊 POSSIBLE CONTEXT VALUES FOR : $M_NAME_FINAL"
                echo "========================================================================================"
                printf "   Model Size:           %.2f GB\n" "$MODEL_SIZE_GB"
                printf "   Optimal GPU Offload:  %s/%s Layers (RAM spill: %s GB)\n" "${NGL_SPILL:-0}" "${TOTAL_LAYERS:-0}" "${RAM_USED_BY_MODEL:-0.00}"
                printf "   Active VRAM Overhead: %s GB\n" "$VRAM_OVERHEAD_GB"
                printf "   Active RAM Overhead:  %s GB\n" "$RAM_OVERHEAD_GB"
                echo "────────────────────────────────────────────────────────────────────────────────────────"
                printf "%-5s | %-32s | %-15s | %-15s\n" "IDX" "STRATEGY PROFILE" "MAX CONTEXT" "KV CACHE TYPE"
                echo "────────────────────────────────────────────────────────────────────────────────────────"
                echo "🌐 Pure VRAM Execution (KV Cache hits GPU VRAM):"
                printf "%-5s | %-32s | %-15s | %-15s\n" "1)" "Pure VRAM (100% GPU)" "${CTX_V_F16:-0}" "F16"
                printf "%-5s | %-32s | %-15s | %-15s\n" "2)" "Pure VRAM (100% GPU)" "${CTX_V_Q8:-0}" "Q8_0"
                printf "%-5s | %-32s | %-15s | %-15s\n" "3)" "Pure VRAM (100% GPU)" "${CTX_V_Q4:-0}" "Q4_0"
                echo "────────────────────────────────────────────────────────────────────────────────────────"
                echo "⚙  Offloading to RAM (KV Cache hits System RAM):"
                printf "%-5s | %-32s | %-15s | %-15s\n" "4)" "RAM Offload (-ngl ${NGL_SPILL:-0})" "$DISP_R_F16" "F16"
                printf "%-5s | %-32s | %-15s | %-15s\n" "5)" "RAM Offload (-ngl ${NGL_SPILL:-0})" "$DISP_R_Q8" "Q8_0"
                printf "%-5s | %-32s | %-15s | %-15s\n" "6)" "RAM Offload (-ngl ${NGL_SPILL:-0})" "$DISP_R_Q4" "Q4_0"
                echo "────────────────────────────────────────────────────────────────────────────────────────"
             
                printf "Select context size or 0 to cancel: "
                read -r strategy_choice
             
                if [ "$strategy_choice" = "4" ] && [ "${CTX_R_F16:-0}" -eq 0 ] 2>/dev/null; then echo "❌ Selection blocked: Will cause Out-Of-Memory crash."; strategy_choice=0; fi
                if [ "$strategy_choice" = "5" ] && [ "${CTX_R_Q8:-0}" -eq 0 ] 2>/dev/null; then echo "❌ Selection blocked: Will cause Out-Of-Memory crash."; strategy_choice=0; fi
                if [ "$strategy_choice" = "6" ] && [ "${CTX_R_Q4:-0}" -eq 0 ] 2>/dev/null; then echo "❌ Selection blocked: Will cause Out-Of-Memory crash."; strategy_choice=0; fi

                if [ "$strategy_choice" != "0" ] && [ -n "$strategy_choice" ]; then
                     case "$strategy_choice" in
                         1) ngl_val="$TOTAL_LAYERS"; q_type="f16";  kqv_offload="true";  ctx_chosen="$CTX_V_F16" ;;
                         2) ngl_val="$TOTAL_LAYERS"; q_type="q8_0"; kqv_offload="true";  ctx_chosen="$CTX_V_Q8" ;;
                         3) ngl_val="$TOTAL_LAYERS"; q_type="q4_0"; kqv_offload="true";  ctx_chosen="$CTX_V_Q4" ;;
                         4) ngl_val="$NGL_SPILL";     q_type="f16";  kqv_offload="false"; ctx_chosen="$CTX_R_F16" ;;
                         5) ngl_val="$NGL_SPILL";     q_type="q8_0"; kqv_offload="false"; ctx_chosen="$CTX_R_Q8" ;;
                         6) ngl_val="$NGL_SPILL";     q_type="q4_0"; kqv_offload="false"; ctx_chosen="$CTX_R_Q4" ;;
                         *) ngl_val="$NGL_SPILL";     q_type="f16";  kqv_offload="false"; ctx_chosen="$CTX_R_F16" ;;
                     esac

                     host_final="${CONFIG_SERVER_HOST:-127.0.0.1}"
                     port_final="${CONFIG_SERVER_PORT:-8080}"

                     echo "🚀 Starting llama server..."
                     "$SCRIPTS_DIR/run-llama_server.sh" \
                         "$SELECTED_MODEL_PATH" \
                         "$ngl_val" \
                         "$q_type" \
                         "$kqv_offload" \
                         "$host_final" \
                         "$port_final" \
                         "$ctx_chosen"
                fi
            else
                echo "❌ Invalid context choice."
            fi
        fi
        rm -f "$SORTED_LIST_FILE"
        ;;

    2)
        if [ -f "$SCRIPTS_DIR/config-params.sh" ]; then
            "$SCRIPTS_DIR/config-params.sh"
        else
            "$HOME/scripts/config-params.sh"
        fi
     
        # Live reload newly stored configuration metrics back into context space
        unset CONFIG_SCRIPTS_PATH CONFIG_MODELS_PATH CONFIG_WORKSPACE_PATH
        unset CONFIG_AIDER_PATH CONFIG_CLAUDE_PATH CONFIG_SERVER_HOST CONFIG_SERVER_PORT
        unset CONFIG_LLAMA_SERVER_BIN CONFIG_DUMMY_UVM_HOOK CONFIG_VRAM_OVERHEAD CONFIG_RAM_OVERHEAD
        unset CONFIG_CONTEXT_MULTIPLIER CONFIG_DEBUG
     
        if [ -f "$CONFIG_FILE" ]; then
            . "$CONFIG_FILE"
            SCRIPTS_DIR="$CONFIG_SCRIPTS_PATH"
            MODELS_DIR="$CONFIG_MODELS_PATH"
            WORKSPACE_DIR="$CONFIG_WORKSPACE_PATH"
            AIDER_BIN_PATH="$CONFIG_AIDER_PATH"
            CLAUDE_BIN_PATH="$CONFIG_CLAUDE_PATH"
            SERVER_HOST_VAL="$CONFIG_SERVER_HOST"
            SERVER_PORT_VAL="$CONFIG_SERVER_PORT"
            LLAMA_SERVER_BIN="$CONFIG_LLAMA_SERVER_BIN"
            DUMMY_UVM_HOOK="$CONFIG_DUMMY_UVM_HOOK"
            VRAM_OVERHEAD_GB="${CONFIG_VRAM_OVERHEAD:-2}"
            RAM_OVERHEAD_GB="${CONFIG_RAM_OVERHEAD:-3.0}"
            CONTEXT_MULTIPLIER="${CONFIG_CONTEXT_MULTIPLIER:-1.5}"
            DEBUG_MODE="${CONFIG_DEBUG:-false}"
        fi
        ;;

    0|*)
        echo "──────────────────────────────────────────────────────"
        echo "🛑 Exiting Dashboard. Initiating deep hardware cleanup..."
        echo "──────────────────────────────────────────────────────"
     
        TARGET_PORT="${SERVER_PORT_VAL:-8080}"
        PID_FILE="/tmp/llama-server.pid"
     
        # 🎯 STRATEGI 1: Tvinga bort processen via nätverksporten med fuser (Mest stensäkert på FreeBSD)
        if command -v fuser >/dev/null 2>&1; then
            echo "🔥 Enforcing fuser port-lock purge on tcp/$TARGET_PORT..."
            # -k = kill, -s = silent/forced signal, -t = tcp port target
            fuser -k -s -t "$TARGET_PORT" 2>/dev/null || true
            sleep 1
        fi

        # 🎯 STRATEGI 2: Sekventiell rensning via den sparade PID-filen
        if [ -f "$PID_FILE" ]; then
            ACTIVE_PID=$(cat "$PID_FILE" | tr -d '\n\r ')
            if [ -n "$ACTIVE_PID" ] && [ "$ACTIVE_PID" -gt 0 ] 2>/dev/null; then
                echo "⏳ Sending termination vector to PID: $ACTIVE_PID..."
                kill -9 "$ACTIVE_PID" 2>/dev/null || true
            fi
            rm -f "$PID_FILE"
        fi
     
        # 🎯 STRATEGI 3: Jaga rätt på dolda Linuxulator-trådar via full strängmatchning
        # (pgrep -f hittar processer inuti /compat/linux/ även om PID har skiftat)
        LLAMA_PIDS=$(pgrep -f "llama-server" 2>/dev/null)
        if [ -n "$LLAMA_PIDS" ]; then
            echo "⏳ Purging orphaned Linuxulator process nodes..."
            echo "$LLAMA_PIDS" | while read -r lp; do
                if [ -n "$lp" ]; then
                    kill -9 "$lp" 2>/dev/null || true
                fi
            done
        fi
     
        # 🎯 STRATEGI 4: Sista universella spärren - Hård killall på binärnamnet
        killall -9 llama-server 2>/dev/null || true
     
        echo "✅ Deep cleanup complete. All VRAM and network ports released."
        echo "Goodbye."
        exit 0
        ;;

esac

done
sh:
#!/bin/sh

# ==============================================================================
# ⚙️ GConfiguration parameters
# ==============================================================================
CONFIG_FILE="$HOME/.git-analyzer.conf"

# Load existing configuration profile defaults if available
if [ -f "$CONFIG_FILE" ]; then
    . "$CONFIG_FILE"
fi

echo "======================================================"
echo "⚙️  APARAMETER CONFIGURATION MANAGER"
echo "======================================================"

# 1. Scripts directory path
printf "Enter scripts folder path [%s]: " "${CONFIG_SCRIPTS_PATH:-$HOME/scripts}"
read -r input_scripts
SCRIPTS_PATH="${input_scripts:-${CONFIG_SCRIPTS_PATH:-$HOME/scripts}}"

# 2. AI models path
printf "Enter models folder path [%s]: " "${CONFIG_MODELS_PATH:-$HOME/AI-models}"
read -r input_models
MODELS_PATH="${input_models:-${CONFIG_MODELS_PATH:-$HOME/AI-models}}"

# 3. Active workspace path
printf "Enter working space path [%s]: " "${CONFIG_WORKSPACE_PATH:-$HOME/workspace}"
read -r input_workspace
WORKSPACE_PATH="${input_workspace:-${CONFIG_WORKSPACE_PATH:-$HOME/workspace}}"

# 4. Aider client binary path
printf "Enter Aider binary path [%s]: " "${CONFIG_AIDER_PATH:-/usr/local/bin/aider}"
read -r input_aider
AIDER_PATH="${input_aider:-${CONFIG_AIDER_PATH:-/usr/local/bin/aider}}"

# 5. Claude client binary path
printf "Enter Claude-code binary path [%s]: " "${CONFIG_CLAUDE_PATH:-/usr/local/bin/claude}"
read -r input_claude
CLAUDE_PATH="${input_claude:-${CONFIG_CLAUDE_PATH:-/usr/local/bin/claude}}"

# 6. Network host ip
printf "Enter Llama server hostname/IP [%s]: " "${CONFIG_SERVER_HOST:-127.0.0.1}"
read -r input_host
SERVER_HOST="${input_host:-${CONFIG_SERVER_HOST:-127.0.0.1}}"

# 7. Network port
printf "Enter Llama server port [%s]: " "${CONFIG_SERVER_PORT:-8080}"
read -r input_port
SERVER_PORT="${input_port:-${CONFIG_SERVER_PORT:-8080}}"

# 8. Llama server executable path
printf "Enter Llama server executable binary path [%s]: " "${CONFIG_LLAMA_SERVER_BIN:-/usr/local/bin/llama-server}"
read -r input_bin
LLAMA_SERVER_BIN_VAL="${input_bin:-${CONFIG_LLAMA_SERVER_BIN:-/usr/local/bin/llama-server}}"

# 9. Dummy-uvm-hook path
printf "Use Dummy UVM Hook? (true/false) [%s]: " "${CONFIG_DUMMY_UVM_HOOK:-false}"
read -r input_uvm
DUMMY_UVM_HOOK_VAL="${input_uvm:-${CONFIG_DUMMY_UVM_HOOK:-false}}"

# 10. VRAM overhead
printf "Enter active VRAM overhead in GB [%s]: " "${CONFIG_VRAM_OVERHEAD:-1.0}"
read -r input_vram
VRAM_OVERHEAD_VAL="${input_vram:-${CONFIG_VRAM_OVERHEAD:-1.0}}"

# 11. RAM overhead
printf "Enter active RAM overhead in GB [%s]: " "${CONFIG_RAM_OVERHEAD:-2.0}"
read -r input_ram
RAM_OVERHEAD_VAL="${input_ram:-${CONFIG_RAM_OVERHEAD:-2.0}}"

# 12. Context windows scaling factor (X)
printf "Context multiplier (X * trained context, e.g. 1.0, 1.5) [%s]: " "${CONFIG_CONTEXT_MULTIPLIER:-1.5}"
read -r input_multiplier
CONTEXT_MULTIPLIER_VAL="${input_multiplier:-${CONFIG_CONTEXT_MULTIPLIER:-1.5}}"

# 13. Debug printout toggle
printf "Enable verbose debug outputs? (true/false) [%s]: " "${CONFIG_DEBUG:-false}"
read -r input_debug
DEBUG_VAL="${input_debug:-${CONFIG_DEBUG:-false}}"

# Save entire normalized parameter manifest automically back into your config file location
cat << EOF > "$CONFIG_FILE"
CONFIG_SCRIPTS_PATH="$SCRIPTS_PATH"
CONFIG_MODELS_PATH="$MODELS_PATH"
CONFIG_WORKSPACE_PATH="$WORKSPACE_PATH"
CONFIG_AIDER_PATH="$AIDER_PATH"
CONFIG_CLAUDE_PATH="$CLAUDE_PATH"
CONFIG_SERVER_HOST="$SERVER_HOST"
CONFIG_SERVER_PORT="$SERVER_PORT"
CONFIG_LLAMA_SERVER_BIN="$LLAMA_SERVER_BIN_VAL"
CONFIG_DUMMY_UVM_HOOK="$DUMMY_UVM_HOOK_VAL"
CONFIG_VRAM_OVERHEAD="$VRAM_OVERHEAD_VAL"
CONFIG_RAM_OVERHEAD="$RAM_OVERHEAD_VAL"
CONFIG_CONTEXT_MULTIPLIER="$CONTEXT_MULTIPLIER_VAL"
CONFIG_DEBUG="$DEBUG_VAL"
EOF

echo "──────────────────────────────────────────────────────"
echo "✅ Parameter configuration profile saved securely to: $CONFIG_FILE"
echo "──────────────────────────────────────────────────────"

/grandpa
 
Last edited:
sh:
#!/bin/sh

# ==============================================================================
# 🖥️  FREEBSD NATIVE HARDWARE TELEMETRY & DISCOVERY LAYER
# ==============================================================================

echo "──────────────────────────────────────────────────────"
echo "🖥️  FCheck VRAM and RAM"
echo "──────────────────────────────────────────────────────"

# 1. DISCOVER SYSTEM RAM (Store and export precise raw bytes)
PHYSMEM_BYTES=$(sysctl -n hw.physmem 2>/dev/null)

if [ -n "$PHYSMEM_BYTES" ]; then
    # Maintain exact raw bytes for calculation scripts while displaying clean GB
    TOTAL_RAM_GB=$(awk -v b="$PHYSMEM_BYTES" 'BEGIN {printf "%.2f", b / 1024 / 1024 / 1024}')
    echo "🧠 System RAM Status:"
    echo "   • Total Physical RAM : $TOTAL_RAM_GB GB ($PHYSMEM_BYTES Bytes)"
else
    echo "❌ Error: Unable to query physical system memory via sysctl."
    PHYSMEM_BYTES=0
    TOTAL_RAM_GB="0.00"
fi

echo "──────────────────────────────────────────────────────"

# 2. DISCOVER NVIDIA GRAPHICS VRAM
echo "🎮 NVIDIA Graphics VRAM Status:"

if [ ! -c /dev/nvidia0 ]; then
    echo "   ❌ Error: NVIDIA hardware device node (/dev/nvidia0) is missing!"
    TOTAL_VRAM_GB="0.00"
    VRAM_BYTES=0
else
    NV_DRIVER_VERSION=$(sysctl -n hw.nvidia.version 2>/dev/null)
    if [ -n "$NV_DRIVER_VERSION" ]; then
        echo "   • Kernel Driver      : NVIDIA UNIX x86_64 ($NV_DRIVER_VERSION)"
    fi

    GPU_MODEL=$(sysctl -n hw.nvidia.0.description 2>/dev/null)
    if [ -n "$GPU_MODEL" ]; then
        echo "   • Hardware Model     : $GPU_MODEL"
    fi

    # Primary discovery method via pciconf BAR registry mapping
    VRAM_BYTES=$(pciconf -lv | grep -A 4 "vgapci" | grep -i "nvidia" -A 4 2>/dev/null | grep "bar.*vram" | sed -E 's/.*vram[[:space:]]*([0-9]+).*/\1/')
 
    # Secondary discovery method fallback via raw sysctl properties
    if [ -z "$VRAM_BYTES" ]; then
        VRAM_MB=$(sysctl -a 2>/dev/null | grep "nvidia" | grep -i "vram" | head -n 1 | awk '{print $NF}' | sed 's/[^0-9]//g')
        [ -n "$VRAM_MB" ] && [ "$VRAM_MB" -gt 0 ] 2>/dev/null && VRAM_BYTES=$((VRAM_MB * 1024 * 1024))
    fi

    # Tertiary discovery method fallback via native FreeBSD nvidia-smi control query interface
    if [ -z "$VRAM_BYTES" ] || [ "$VRAM_BYTES" -eq 0 ]; then
        FREEBSD_SMI="/usr/local/bin/nvidia-smi"
        if [ ! -x "$FREEBSD_SMI" ]; then
            FREEBSD_SMI=$(command -v nvidia-smi 2>/dev/null || echo "")
        fi
  
        if [ -n "$FREEBSD_SMI" ] && [ -x "$FREEBSD_SMI" ]; then
            VRAM_SMI_MB=$("$FREEBSD_SMI" --query-gpu=memory.total --format=csv,noheader,nounits 2>/dev/null | tr -d '\n\r ')
            [ -n "$VRAM_SMI_MB" ] && [ "$VRAM_SMI_MB" -gt 0 ] 2>/dev/null && VRAM_BYTES=$((VRAM_SMI_MB * 1024 * 1024))
        fi
    fi

    # Dynamic protection gate: Prevent empty or zero variables from polluting the calculations
    if [ -z "$VRAM_BYTES" ] || [ "$VRAM_BYTES" -eq 0 ]; then
        echo "⚠️  Telemetry Warning: Dynamic VRAM discovery failed. Check driver load status!" >&2
        VRAM_BYTES=0
    fi

    TOTAL_VRAM_GB=$(awk -v b="$VRAM_BYTES" 'BEGIN {printf "%.2f", b / 1024 / 1024 / 1024}')
    echo "   • Dedicated VRAM     : $TOTAL_VRAM_GB GB ($VRAM_BYTES Bytes)"
fi

echo "──────────────────────────────────────────────────────"
echo "✅ Discovery phase complete."
echo "──────────────────────────────────────────────────────"

# Export the precise binary byte values directly into memory registry context
export INVENTORIED_RAM="$PHYSMEM_BYTES"
export INVENTORIED_VRAM="$VRAM_BYTES"
sh:
#!/bin/sh

# Exit immediately with a zeroed string if no model file is provided
if [ -z "$1" ] || [ ! -f "$1" ]; then
    echo "0|0|0"
    exit 1
fi

GGUF_FILE="$1"
CONFIG_FILE="$HOME/.git-analyzer.conf"

# Establish baseline fallback defaults
VRAM_OVERHEAD_GB=0
CONTEXT_MULTIPLIER=1.5
SCRIPTS_DIR="$HOME/scripts"

# Load global configuration directives dynamically if present
if [ -f "$CONFIG_FILE" ]; then
    . "$CONFIG_FILE"
    VRAM_OVERHEAD_GB="${CONFIG_VRAM_OVERHEAD:-0}"
    CONTEXT_MULTIPLIER="${CONFIG_CONTEXT_MULTIPLIER:-1.5}"
    SCRIPTS_DIR="${CONFIG_SCRIPTS_PATH:-$HOME/scripts}"
fi

# Fetch VRAM and RAM via the script
HW_OUTPUT=$("$SCRIPTS_DIR/check-vram_ram.sh" 2>/dev/null)
TOTAL_VRAM_BYTES=$(echo "$HW_OUTPUT" | grep -i "Dedicated VRAM" | awk '{print $NF}' | sed 's/[^0-9]//g')
[ -z "$TOTAL_VRAM_BYTES" ] && TOTAL_VRAM_BYTES=12884901888

# Compute net available VRAM bytes after subtracting user-configured safety overhead
AVAILABLE_VRAM_BYTES=$(awk -v t="$TOTAL_VRAM_BYTES" -v o="$VRAM_OVERHEAD_GB" 'BEGIN {print t - (o * 1024 * 1024 * 1024)}')
FILE_SIZE_BYTES=$(stat -f %z "$GGUF_FILE" 2>/dev/null || stat -c %s "$GGUF_FILE")

# Parse model layers and attention metrics from the GGUF header dictionary
MODEL_STATS=$("$SCRIPTS_DIR/get-model_gguf_data.py" "$GGUF_FILE" 2>/dev/null)
TOTAL_LAYERS=$(echo "$MODEL_STATS" | grep "Number of Layers:" | awk '{print $NF}')
HEAD_DIM=$(echo "$MODEL_STATS" | grep "Head Dimension:" | awk '{print $NF}')
KV_HEADS=$(echo "$MODEL_STATS" | grep "Head Count KV:" | awk '{print $NF}')
TRAINED_CTX=$(echo "$MODEL_STATS" | grep "Trained Context:" | awk '{print $NF}')

# Revert to fallback ceilings if metadata registers as unknown
[ -z "$TRAINED_CTX" ] || [ "$TRAINED_CTX" = "Unknown" ] && TRAINED_CTX=32768

# Calculate the dynamic ceiling barrier (X * Trained Context)
MAX_ALLOWED_CTX=$(awk -v tc="$TRAINED_CTX" -v m="$CONTEXT_MULTIPLIER" 'BEGIN {printf "%.0f", tc * m}')

# Validate critical model architecture metrics before executing division loops
if [ -z "$TOTAL_LAYERS" ] || [ "$TOTAL_LAYERS" = "Unknown" ] || [ "$TOTAL_LAYERS" -eq 0 ] 2>/dev/null; then
    echo "0|0|0"
    exit 1
fi

BYTES_PER_LAYER=$(awk -v size="$FILE_SIZE_BYTES" -v l="$TOTAL_LAYERS" 'BEGIN {printf "%.0f", size / l}')
NGL_SPILL=$(awk -v vram="$AVAILABLE_VRAM_BYTES" -v bpl="$BYTES_PER_LAYER" -v l="$TOTAL_LAYERS" 'BEGIN {ngl = int(vram / bpl); print (ngl < 0 ? 0 : (ngl > l ? l : ngl))}')

# Evaluate Pure VRAM pipeline options only if 100% of network layers offload to GPU memory
if [ "$NGL_SPILL" = "$TOTAL_LAYERS" ]; then
    FREE_VRAM_BYTES=$(awk -v vram="$AVAILABLE_VRAM_BYTES" -v mod="$FILE_SIZE_BYTES" 'BEGIN {print vram - mod}')
    [ "$FREE_VRAM_BYTES" -le 0 ] && FREE_VRAM_BYTES=104857600

    # Calculate absolute physical maximum token ranges based on free hardware sectors
    CTX_V_F16=$(awk -v free="$FREE_VRAM_BYTES" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 2; print (denom > 0 ? int(free / denom) : 0)}')
    CTX_V_Q8=$(awk -v free="$FREE_VRAM_BYTES" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 1.0625; print (denom > 0 ? int(free / denom) : 0)}')
    CTX_V_Q4=$(awk -v free="$FREE_VRAM_BYTES" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 0.5625; print (denom > 0 ? int(free / denom) : 0)}')

    # Clamp the computational parameters directly against the dynamic multiplier ceiling
    [ "$CTX_V_F16" -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_V_F16="$MAX_ALLOWED_CTX"
    [ "$CTX_V_Q8"  -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_V_Q8="$MAX_ALLOWED_CTX"
    [ "$CTX_V_Q4"  -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_V_Q4="$MAX_ALLOWED_CTX"
else
    CTX_V_F16=0; CTX_V_Q8=0; CTX_V_Q4=0
fi

# Print out the final values
echo "$CTX_V_F16|$CTX_V_Q8|$CTX_V_Q4"
sh:
#!/bin/sh

# Exit immediately with a zeroed string if no model file is provided
if [ -z "$1" ] || [ ! -f "$1" ]; then
echo "0|0|0|0|0|0"
exit 1
fi

GGUF_FILE="$1"
CONFIG_FILE="$HOME/.git-analyzer.conf"

# Establish baseline fallback defaults
VRAM_OVERHEAD_GB=0
RAM_OVERHEAD_GB=0
CONTEXT_MULTIPLIER=1.5
SCRIPTS_DIR="$HOME/scripts"

# Load global configuration directives dynamically if present
if [ -f "$CONFIG_FILE" ]; then
. "$CONFIG_FILE"
VRAM_OVERHEAD_GB="${CONFIG_VRAM_OVERHEAD:-0}"
RAM_OVERHEAD_GB="${CONFIG_RAM_OVERHEAD:-0}"
CONTEXT_MULTIPLIER="${CONFIG_CONTEXT_MULTIPLIER:-1.5}"
SCRIPTS_DIR="${CONFIG_SCRIPTS_PATH:-$HOME/scripts}"
fi

# Fetch VRAM and RAM from script
HW_OUTPUT=$("$SCRIPTS_DIR/check-vram_ram.sh" 2>/dev/null)
TOTAL_RAM_BYTES=$(echo "$HW_OUTPUT" | grep -i "Total Physical RAM" | awk '{print $NF}' | sed 's/[^0-9]//g')
TOTAL_VRAM_BYTES=$(echo "$HW_OUTPUT" | grep -i "Dedicated VRAM" | awk '{print $NF}' | sed 's/[^0-9]//g')

[ -z "$TOTAL_RAM_BYTES" ] && TOTAL_RAM_BYTES=17179869184
[ -z "$TOTAL_VRAM_BYTES" ] && TOTAL_VRAM_BYTES=12884901888

# Compute net available VRAM bytes after subtracting user-configured safety overhead
AVAILABLE_VRAM_BYTES=$(awk -v t="$TOTAL_VRAM_BYTES" -v o="$VRAM_OVERHEAD_GB" 'BEGIN {print t - (o * 1024 * 1024 * 1024)}')
FILE_SIZE_BYTES=$(stat -f %z "$GGUF_FILE" 2>/dev/null || stat -c %s "$GGUF_FILE")

# Parse model layers and attention metrics from the GGUF header dictionary
MODEL_STATS=$("$SCRIPTS_DIR/get-model_gguf_data.py" "$GGUF_FILE" 2>/dev/null)
TOTAL_LAYERS=$(echo "$MODEL_STATS" | grep "Number of Layers:" | awk '{print $NF}')
HEAD_DIM=$(echo "$MODEL_STATS" | grep "Head Dimension:" | awk '{print $NF}')
KV_HEADS=$(echo "$MODEL_STATS" | grep "Head Count KV:" | awk '{print $NF}')
TRAINED_CTX=$(echo "$MODEL_STATS" | grep "Trained Context:" | awk '{print $NF}')

# Revert to fallback ceilings if metadata registers as unknown
[ -z "$TRAINED_CTX" ] || [ "$TRAINED_CTX" = "Unknown" ] && TRAINED_CTX=32768
MAX_ALLOWED_CTX=$(awk -v tc="$TRAINED_CTX" -v m="$CONTEXT_MULTIPLIER" 'BEGIN {printf "%.0f", tc * m}')

# Validate critical model architecture metrics before executing division loops
if [ -z "$TOTAL_LAYERS" ] || [ "$TOTAL_LAYERS" = "Unknown" ] || [ "$TOTAL_LAYERS" -eq 0 ] 2>/dev/null; then
echo "0|0|0|0|0|0"
exit 1
fi

BYTES_PER_LAYER=$(awk -v size="$FILE_SIZE_BYTES" -v l="$TOTAL_LAYERS" 'BEGIN {printf "%.0f", size / l}')
NGL_SPILL=$(awk -v vram="$AVAILABLE_VRAM_BYTES" -v bpl="$BYTES_PER_LAYER" -v l="$TOTAL_LAYERS" 'BEGIN {ngl = int(vram / bpl); print (ngl < 0 ? 0 : (ngl > l ? l : ngl))}')
VRAM_USED_BY_MODEL=$(awk -v ngl="$NGL_SPILL" -v bpl="$BYTES_PER_LAYER" 'BEGIN {printf "%.0f", ngl * bpl}')

# Force any minor precision/rounding differentials under 1MB to exactly 0 bytes
RAM_USED_BY_MODEL=$(awk -v tot="$FILE_SIZE_BYTES" -v vram="$VRAM_USED_BY_MODEL" 'BEGIN {diff = tot - vram; if (diff < 1048576) diff = 0; print diff}')

ACTUAL_FREE_RAM=$(awk -v tot="$TOTAL_RAM_BYTES" -v used="$RAM_USED_BY_MODEL" 'BEGIN {print tot - used}')
FREE_RAM_AVAILABLE=$(awk -v free="$ACTUAL_FREE_RAM" -v oh="$RAM_OVERHEAD_GB" 'BEGIN {margin = oh * 1024 * 1024 * 1024; val = free - margin; print (val < 0 ? 0 : val)}')

# Calculate absolute physical maximum tokens allocatable within net available system RAM
CTX_R_F16=$(awk -v free="$FREE_RAM_AVAILABLE" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 2; print (denom > 0 ? int(free / denom) : 0)}')
CTX_R_Q8=$(awk -v free="$FREE_RAM_AVAILABLE" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 1.0625; print (denom > 0 ? int(free / denom) : 0)}')
CTX_R_Q4=$(awk -v free="$FREE_RAM_AVAILABLE" -v l="$TOTAL_LAYERS" -v kh="$KV_HEADS" -v hd="$HEAD_DIM" 'BEGIN {denom = 2 * l * kh * hd * 0.5625; print (denom > 0 ? int(free / denom) : 0)}')

# Clamp the offloading context variables safely against the dynamic multiplier ceiling
[ "$CTX_R_F16" -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_R_F16="$MAX_ALLOWED_CTX"
[ "$CTX_R_Q8" -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_R_Q8="$MAX_ALLOWED_CTX"
[ "$CTX_R_Q4" -gt "$MAX_ALLOWED_CTX" ] 2>/dev/null && CTX_R_Q4="$MAX_ALLOWED_CTX"

RAM_SPILL_GB=$(awk -v b="$RAM_USED_BY_MODEL" 'BEGIN {print b / 1024 / 1024 / 1024}')

# Print out the complete system offloading values
echo "$NGL_SPILL|$RAM_SPILL_GB|$CTX_R_F16|$CTX_R_Q8|$CTX_R_Q4|$TOTAL_LAYERS"


/grandpa
 
Last edited:
Python:
#!/usr/local/bin/python3.12
import sys
import struct
import os

def read_gguf_metadata(filepath):
    data = {
        "architecture": "Unknown",
        "block_count": "Unknown",
        "embedding_length": "Unknown",
        "head_count": "Unknown",
        "head_count_kv": "Unknown",
        "context_length": "Unknown",
        "file_size_bytes": "Unknown"
    }
  
    data["file_size_bytes"] = os.path.getsize(filepath)
  
    with open(filepath, "rb") as f:
        # 1. Verify Magic Header
        magic = f.read(4)
        if magic != b"GGUF":
            return data

        # 2. Read GGUF Version
        version = struct.unpack("<I", f.read(4))[0]
        if version not in (2, 3):
            return data
      
        # 3. Read total tensor count and metadata KV count
        tensor_count = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
        kv_count = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]

        # 4. Iterate over key-value metadata dictionaries
        for i in range(kv_count):
            key_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
            key = f.read(key_len).decode("utf-8", errors="ignore")
            val_type = struct.unpack("<I", f.read(4))[0]
              
            def parse_value(t):
                if t in (0, 1, 7):
                    return struct.unpack("<B" if t==0 else "<b" if t==1 else "?", f.read(1))[0]
                elif t in (2, 3):
                    return struct.unpack("<H" if t==2 else "<h", f.read(2))[0]
                elif t in (4, 5, 6):
                    fmt = "<I" if t==4 else "<i" if t==5 else "<f"
                    return struct.unpack(fmt, f.read(4))[0]
                elif t == 8:
                    s_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
                    return f.read(s_len).decode("utf-8", errors="ignore")
                elif t == 9:
                    arr_type = struct.unpack("<I", f.read(4))[0]
                    arr_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
                    res_list = []
                    for _ in range(arr_len):
                        res_list.append(parse_value(arr_type))
                    return res_list
                elif t in (10, 11, 12, 13, 14):
                    fmt = "<Q" if t in (10, 13) else "<q" if t in (11, 14) else "<d"
                    return struct.unpack(fmt, f.read(8))[0]
                return None
  
            res = parse_value(val_type)

            # 5. Store parsed parameter targets
            if key == "general.architecture":
                data["architecture"] = res
            elif key.endswith(".block_count"):
                data["block_count"] = res
            elif key.endswith(".embedding_length"):
                data["embedding_length"] = res
            elif key.endswith(".attention.head_count"):
                data["head_count"] = res
            elif key.endswith(".attention.head_count_kv"):
                data["head_count_kv"] = res
            elif key.endswith(".context_length"):
                data["context_length"] = res

            if key.startswith("tensor."):
                break
              
    return data

if __name__ == "__main__":
    if len(sys.argv) > 1:
        metadata = read_gguf_metadata(sys.argv[1])
      
        # Fallback for standard MHA (Multi-Head Attention) if KV heads are not explicitly defined
        if metadata["head_count_kv"] == "Unknown" and metadata["head_count"] != "Unknown":
            metadata["head_count_kv"] = metadata["head_count"]

        # FIX: Extract inner values correctly if returned inside structured lists
        if isinstance(metadata["architecture"], list) and len(metadata["architecture"]) > 0:
            metadata["architecture"] = metadata["architecture"][0]

        if isinstance(metadata["head_count_kv"], list) and len(metadata["head_count_kv"]) > 0:
            metadata["head_count_kv"] = metadata["head_count_kv"][0]
          
        if isinstance(metadata["head_count"], list) and len(metadata["head_count"]) > 0:
            metadata["head_count"] = metadata["head_count"][0]
          
        if isinstance(metadata["block_count"], list) and len(metadata["block_count"]) > 0:
            metadata["block_count"] = metadata["block_count"][0]
          
        if isinstance(metadata["embedding_length"], list) and len(metadata["embedding_length"]) > 0:
            metadata["embedding_length"] = metadata["embedding_length"][0]

        if isinstance(metadata["context_length"], list) and len(metadata["context_length"]) > 0:
            metadata["context_length"] = metadata["context_length"][0]

        # Extract dimensions based on newly matched integers
        layers = metadata["block_count"]
        hidden_size = metadata["embedding_length"]
      
        # Calculate Head Dimension (Hidden Size / Head Count)
        if metadata["embedding_length"] != "Unknown" and metadata["head_count"] != "Unknown":
            head_dim = metadata["embedding_length"] // metadata["head_count"]
        else:
            head_dim = "Unknown"

        # Calculate GQA/MQA scaling divisor factor (Query Heads / KV Heads)
        if metadata["head_count"] != "Unknown" and metadata["head_count_kv"] != "Unknown":
            gqa_factor = metadata["head_count"] // metadata["head_count_kv"]
        else:
            gqa_factor = "Unknown"

        print(f"Architecture:      {metadata['architecture']}")
        print(f"Block Count:       {metadata['block_count']}")
        print(f"Embedding Length:  {metadata['embedding_length']}")
        print(f"Head Count:        {metadata['head_count']}")
        print(f"Head Count KV:     {metadata['head_count_kv']}")
        print(f"File Size (Bytes): {metadata['file_size_bytes']}")
        print(f"Number of Layers:  {layers}")
        print(f"Hidden Size:       {hidden_size}")
        print(f"Head Dimension:    {head_dim}")
        print(f"GQA/MQA Factor:    1:{gqa_factor}" if gqa_factor != "Unknown" else "GQA/MQA Factor:    Unknown")
        print(f"Trained Context:   {metadata['context_length']}")
    else:
        print("Usage: ./get-model_gguf_data.py <path_to_model.gguf>")
sh:
#!/bin/sh

# ==============================================================================
# 🤖 CLAUDE CODE CLI RUNNER (Modular & Optimized Local Edition)
# ==============================================================================
# Purpose: Dynamically maps Anthropic environmental layers to perfectly match
#          the hardware and token boundaries calculated by the matrix engine.
#          Utilizes full parameter paths and aggressive context compression.
# ==============================================================================

CONFIG_FILE="$HOME/.git-analyzer.conf"

# Establish baseline fallback defaults
AIDER_TARGET_HOST="127.0.0.1"
SERVER_PORT="8080"
DEBUG_MODE="false"
WORKSPACE_DIR="$HOME/workspace"

# Load global configuration directives dynamically if present
if [ -f "$CONFIG_FILE" ]; then
    . "$CONFIG_FILE"
    AIDER_TARGET_HOST="${CONFIG_SERVER_HOST:-127.0.0.1}"
    SERVER_PORT="${CONFIG_SERVER_PORT:-8080}"
    DEBUG_MODE="${CONFIG_DEBUG:-false}"
    WORKSPACE_DIR="${CONFIG_WORKSPACE_PATH:-$HOME/workspace}"
fi

# Extract optimized runtime parameters forwarded from the execution bridge
SELECTED_MODEL="$1"
SELECTED_CTX="$2"
CLAUDE_BIN_PATH="$3"
TRAINED_CTX="$4"

if [ -z "$TRAINED_CTX" ] || [ "$TRAINED_CTX" = "Unknown" ] 2>/dev/null; then
    TRAINED_CTX="$SELECTED_CTX"
fi

if [ -z "$SELECTED_MODEL" ] || [ -z "$SELECTED_CTX" ] || [ -z "$CLAUDE_BIN_PATH" ]; then
    echo "❌ Execution Error: Missing dynamic context or executable matrix parameters." >&2
    exit 1
fi

# Resolve 0.0.0.0 out to loopback interface for local API routing stability
if [ "$AIDER_TARGET_HOST" = "0.0.0.0" ]; then
    AIDER_TARGET_HOST="127.0.0.1"
fi

# 1. DYNAMICALLY COMPUTE CLAUDE BOUNDARIES BASED ON ENGINE CALCULATIONS
if [ "$SELECTED_CTX" -lt "$TRAINED_CTX" ] 2>/dev/null; then
    SAFE_LIMIT=$SELECTED_CTX
else
    SAFE_LIMIT=$TRAINED_CTX
fi

COMPACT_WINDOW=$(( SAFE_LIMIT * 80 / 100 ))

# ==============================================================================
# 2. ACTIVATED: LOCAL LOCAL-MODEL TUNING & AGENT COMPRESSION
# ==============================================================================
export ANTHROPIC_MODEL="claude-sonnet-5"
export ANTHROPIC_BASE_URL="http://${AIDER_TARGET_HOST}:${SERVER_PORT}"
export ANTHROPIC_DEFAULT_SONNET_MODEL="$SELECTED_MODEL"
export ANTHROPIC_DEFAULT_OPUS_MODEL="claude-sonnet-5"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="claude-sonnet-5"
export ANTHROPIC_AUTH_TOKEN="ollama"
export ANTHROPIC_SMALL_FAST_MODEL="claude-sonnet-5"

# Active execution bounds mapped directly to your dynamic matrix headroom
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="$COMPACT_WINDOW"
export CLAUDE_AUTOCOMPACT_PCT_OVERRIDE="80"
export CLAUDE_CODE_MAX_OUTPUT_TOKENS="8192"

# 🎯 CRITICAL LOCAL-MODEL LOCKS (Prevents channels loops and hallucinations)
export ANTHROPIC_STREAMING_JSON_PATCH="0"
export CLAUDE_CODE_OUTPUT_FORMAT="text"
export CLAUDE_CODE_DISABLE_FAST_MODE="1"
export CLAUDE_CODE_MAX_CHAT_HISTORY_TOKENS="$(( SAFE_LIMIT / 2 ))"
export CLAUDE_CODE_DISABLE_CHANNELS="1"
export CLAUDE_CODE_DISABLE_JSON_SCHEMA_VALIDATION="0"
export MCP_PROTOCOL_NEGOTIATION="legacy"
export MCP_TOOL_TIMEOUT="60000"

# Secure the native FreeBSD Bash path context for terminal sub-invocations
export CLAUDE_CODE_GIT_BASH_PATH="/usr/local/bin/bash"

# Stop silent background telemetry network tasks and update lookups
export CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY="1"
export DISABLE_AUTOUPDATER="1"
export DISABLE_UPDATES="1"

# ==============================================================================
# 🔍 CLAUDE_CODE CLIENT RUNTIME DEBUG LAYER (Toggled Conditionally)
# ==============================================================================
if [ "$DEBUG_MODE" = "true" ]; then
    echo ""
    echo "--- Llama.cpp Configuration Loaded ---"
    echo "Model Selected : $SELECTED_MODEL (Mapped to Sonnet-5 Profile)"
    echo "LAN Endpoint   : $ANTHROPIC_BASE_URL"
    echo "Calculated CTX : $SELECTED_CTX tokens"
    echo "Compact Buffer : $COMPACT_WINDOW tokens (Triggers auto-rolling shift at 80%)"
    echo "Executable     : $CLAUDE_BIN_PATH"
    echo "────────────────────────────────────────────────────────────────────────"
    echo "🔍 RUN-CLAUDE.SH ENVIRONMENT PARAMETERS DEBUG:"
    echo "   • ANTHROPIC_MODEL         : '$ANTHROPIC_MODEL'"
    echo "   • ANTHROPIC_BASE_URL      : '$ANTHROPIC_BASE_URL'"
    echo "   • AUTO_COMPACT_WINDOW     : '$CLAUDE_CODE_AUTO_COMPACT_WINDOW'"
    echo "   • MAX_OUTPUT_TOKENS       : '$CLAUDE_CODE_MAX_OUTPUT_TOKENS'"
    echo "   • MAX_CHAT_HISTORY_TOKENS : '$CLAUDE_CODE_MAX_CHAT_HISTORY_TOKENS'"
    echo "   • DISABLE_FAST_MODE       : '$CLAUDE_CODE_DISABLE_FAST_MODE'"
    echo "   • DISABLE_CHANNELS        : '$CLAUDE_CODE_DISABLE_CHANNELS'"
    echo "   • MCP_NEGOTIATION         : '$MCP_PROTOCOL_NEGOTIATION'"
    echo "   • MCP_TOOL_TIMEOUT        : '$MCP_TOOL_TIMEOUT'"
    echo "   • DISABLE_SCHEMA_VALID    : '$CLAUDE_CODE_DISABLE_JSON_SCHEMA_VALIDATION'"
    echo "────────────────────────────────────────────────────────────────────────"
#    echo "⏸  Debug Profile Loaded. Press [ENTER] to execute Claude Code or [Ctrl+C] to abort..."
#    read -r _
fi

# Move into the assigned development project workspace directory before initiating Claude
if [ -n "$WORKSPACE_DIR" ] && [ -d "$WORKSPACE_DIR" ]; then
    cd "$WORKSPACE_DIR"
fi

echo "🤖 Launching Claude Code with active deep structural tracing..."
echo "────────────────────────────────────────────────────────────────────────"

# 3. Fire up Claude Code utilizing dynamic variable parameters via clean POSIX memory handoff
exec "$CLAUDE_BIN_PATH" \
    --permission-mode bypassPermissions \
    --dangerously-skip-permissions \
    --model claude-sonnet-5 \
    --mcp-config "$HOME/.claude.json"
sh:
#!/bin/sh

# ==============================================================================
# 🚀 START AIDER
# ==============================================================================
CONFIG_FILE="$HOME/.git-analyzer.conf"

# Establish baseline fallback defaults
SCRIPTS_DIR="$HOME/scripts"
AIDER_BIN_EXEC="aider"
AIDER_TARGET_HOST="127.0.0.1"
SERVER_PORT="8080"
WORKSPACE_DIR="$HOME/workspace"
DEBUG_MODE="false"

# Load global configuration directives dynamically if present
if [ -f "$CONFIG_FILE" ]; then
    . "$CONFIG_FILE"
    SCRIPTS_DIR="${CONFIG_SCRIPTS_PATH:-$HOME/scripts}"
    AIDER_BIN_EXEC="${CONFIG_AIDER_PATH:-aider}"
    AIDER_TARGET_HOST="${CONFIG_SERVER_HOST:-127.0.0.1}"
    SERVER_PORT="${CONFIG_SERVER_PORT:-8080}"
    WORKSPACE_DIR="${CONFIG_WORKSPACE_PATH:-$HOME/workspace}"
    DEBUG_MODE="${CONFIG_DEBUG:-false}"
fi

# 1. EXTRACT PASSED RUNTIME ARGUMENTS
TARGET_MODEL_PATH="$1"
SELECTED_MODEL=$(basename "$TARGET_MODEL_PATH")
SELECTED_CTX="$2"

if [ -z "$SELECTED_CTX" ] || [ "$SELECTED_CTX" -eq 0 ] 2>/dev/null; then
    SELECTED_CTX=8192
fi

# Extract model metrics from the GGUF reader
MODEL_STATS=$("$SCRIPTS_DIR/get-model_gguf_data.py" "$TARGET_MODEL_PATH" 2>/dev/null)
TRAINED_CTX=$(echo "$MODEL_STATS" | grep "Embedding Length:" | awk '{print $NF}')
[ -z "$TRAINED_CTX" ] || [ "$TRAINED_CTX" = "Unknown" ] && TRAINED_CTX="$SELECTED_CTX"

# Resolve 0.0.0.0 out to loopback interface
if [ "$AIDER_TARGET_HOST" = "0.0.0.0" ]; then
    AIDER_TARGET_HOST="127.0.0.1"
fi

# 2. MAP CORE API ROUTING AND MODEL SELECTION LAYERS
export OPENAI_API_BASE="http://${AIDER_TARGET_HOST}:${SERVER_PORT}/v1"
export OPENAI_API_KEY="dummy"
export AIDER_MODEL="openai/$SELECTED_MODEL"
export AIDER_MODEL_ARCHITECTURE="${AIDER_MODEL_ARCHITECTURE:-openai}"
export AIDER_BODY_PROMPT="You are a strict FreeBSD C developer. Output the requested code directly inside a code block. No chatter."
export AIDER_TEMPERATURE="0.1"

# ==============================================================================
# 3. APPLY HARD OPTIMIZATION AND MEMORY RESTRICTIONS (SQZ ADAPTIVE CHECK)
# ==============================================================================
if [ "$TRAINED_CTX" -lt 8192 ] 2>/dev/null; then
    export AIDER_MAP_TOKENS="0"
else
    export AIDER_MAP_TOKENS="1024"
fi

if [ -x "$HOME/.cargo/bin/sqz" ]; then
    SQZ_BIN="$HOME/.cargo/bin/sqz"
elif command -v sqz >/dev/null 2>&1; then
    SQZ_BIN="sqz"
else
    SQZ_BIN=""
fi

if [ -n "$SQZ_BIN" ] && [ "$TRAINED_CTX" -ge 8192 ] 2>/dev/null; then
    VIRTUAL_CTX=$(( SELECTED_CTX * 1 ))
    export AIDER_MAX_CONTEXT_WINDOW="$VIRTUAL_CTX"
    USE_SQZ="true"
else
    export AIDER_MAX_CONTEXT_WINDOW="$SELECTED_CTX"
    USE_SQZ="false"
fi

HALF_TRAINED=$(( TRAINED_CTX / 2 ))
export AIDER_MAX_CHAT_HISTORY_TOKENS="$HALF_TRAINED"

# ==============================================================================
# 4. DYNAMIC METADATA GENERATION FOR UNKNOWN LOCAL MODELS
# ==============================================================================
METADATA_FILE="/tmp/aider_model_metadata_${SELECTED_MODEL}.json"
if [ "$TRAINED_CTX" -lt 8192 ] 2>/dev/null; then
    OUT_TOKENS=256
else
    OUT_TOKENS=8192
fi

cat << EOF > "$METADATA_FILE"
{
  "local": {
    "max_tokens": ${AIDER_MAX_CONTEXT_WINDOW},
    "max_input_tokens": ${AIDER_MAX_CONTEXT_WINDOW},
    "max_context_tokens": ${AIDER_MAX_CONTEXT_WINDOW},
    "max_output_tokens": ${OUT_TOKENS},
    "input_cost_per_token": 0.0,
    "output_cost_per_token": 0.0
  }
}
EOF

export AIDER_MODEL_METADATA_FILE="$METADATA_FILE"

# ==============================================================================
# 5. INITIALIZE THE COGNITIVE ROUTER PARSER (Granite Architecture Tuning)
# ==============================================================================
echo "🤖 Initializing Aider Client Session..."

AIDER_TARGET_MODEL="local"

if echo "$SELECTED_MODEL" | grep -iq "granite-4.2"; then
    EDIT_FORMAT="whole"
    echo "📊 Edit Format     : Whole File Overwrite (Optimized for Granite-4.2)"
else
    EDIT_FORMAT="diff"
    echo "📊 Edit Format     : Standard Git Diff"
fi

# ==============================================================================
# 🔍 DEBUG PRINTOUT
# ==============================================================================
if [ "$DEBUG_MODE" = "true" ]; then
    echo "────────────────────────────────────────────────────────────────────────"
    echo "🔍 RUN-AIDER.SH EXECUTION PROFILE:"
    echo "   • EXEC BINARY        : '$AIDER_BIN_EXEC'"
    echo "   • TARGET MODEL       : '$AIDER_TARGET_MODEL' ($SELECTED_MODEL)"
    echo "   • METADATA FILE PATH : '$METADATA_FILE'"
    echo "   • MAP TOKENS VALUE   : '$AIDER_MAP_TOKENS'"
    echo "   • MAX CONTEXT WINDOW : '$AIDER_MAX_CONTEXT_WINDOW'"
    echo "   • MAX HISTORY TOKENS : '$AIDER_MAX_CHAT_HISTORY_TOKENS'"
    echo "   • USE_SQZ STATUS     : '$USE_SQZ'"
    echo "────────────────────────────────────────────────────────────────────────"
    echo "🌐 ACTIVE API & SAMPLING ENVIRONMENTS:"
    echo "   • OPENAI_API_BASE    : '${OPENAI_API_BASE:-not set}'"
    echo "   • OPENAI_API_KEY     : '${OPENAI_API_KEY:-not set}'"
    echo "   • AIDER_MODEL        : '${AIDER_MODEL:-not set}'"
    echo "   • AIDER_TEMPERATURE  : '${AIDER_TEMPERATURE:-not set}'"
    echo "   • PATH               : '$PATH'"
    echo "────────────────────────────────────────────────────────────────────────"
    echo "🚀 FULL AIDER COMMAND EXECUTION STRING:"
    echo "$AIDER_BIN_EXEC \\"
    echo "  --model \"$AIDER_TARGET_MODEL\" \\"
    echo "  --openai-api-base \"$OPENAI_API_BASE\" \\"
    echo "  --model-metadata-file \"$METADATA_FILE\" \\"
    echo "  --edit-format $EDIT_FORMAT \\"
    echo "  --map-tokens \"$AIDER_MAP_TOKENS\" \\"
    echo "  --no-suggest-shell-commands \\"
    echo "  --yes-always"
    echo "────────────────────────────────────────────────────────────────────────"
#    echo "⏸  Debug Profile Loaded. Press [ENTER] to execute Aider or [Ctrl+C] to abort..."
#    read -r _
fi

# Move into the assigned development project workspace directory before initiating Aider
if [ -n "$WORKSPACE_DIR" ] && [ -d "$WORKSPACE_DIR" ]; then
    cd "$WORKSPACE_DIR"
fi

TEST_RUNNER="$SCRIPTS_DIR/test-runner.sh"

# ==============================================================================
# 6. START AIDER
# ==============================================================================
if [ "$USE_SQZ" = "true" ]; then
    echo "📊 Context Layer   : SQZ Intelligence Active [ojuschugh1/sqz]"
    echo "📊 Hardware Buffer : ${SELECTED_CTX} tokens | Virtual Window: ${AIDER_MAX_CONTEXT_WINDOW} tokens"
    echo "────────────────────────────────────────────────────────────────────────"
 
    exec "$AIDER_BIN_EXEC" \
     --model "openai/$AIDER_TARGET_MODEL" \
     --openai-api-base "$OPENAI_API_BASE" \
     --model-metadata-file "$METADATA_FILE" \
     --edit-format "$EDIT_FORMAT" \
     --map-tokens "$AIDER_MAP_TOKENS" \
     --no-suggest-shell-commands \
     --lint-cmd "c: sh -c 'clang -fsyntax-only \"\$@\" > /tmp/aider_lint.log 2>&1; RC=\$?; cat /tmp/aider_lint.log | $SQZ_BIN; exit \$RC'" \
     --auto-test --test-cmd "sh -c '$TEST_RUNNER > /tmp/aider_test.log 2>&1; RC=\$?; cat /tmp/aider_test.log | $SQZ_BIN; cat /tmp/aider_test.log; exit \$RC'" \
     --git \
     --no-show-model-warnings --yes-always
else
    echo "📊 Context Layer   : Native Hardware Direct (SQZ not found)"
    echo "📊 Allocated Buffer: ${AIDER_MAX_CONTEXT_WINDOW} tokens"
    echo "────────────────────────────────────────────────────────────────────────"
 
    exec "$AIDER_BIN_EXEC" \
     --model "openai/$AIDER_TARGET_MODEL" \
     --openai-api-base "$OPENAI_API_BASE" \
     --model-metadata-file "$METADATA_FILE" \
     --edit-format "$EDIT_FORMAT" \
     --map-tokens "$AIDER_MAP_TOKENS" \
     --no-suggest-shell-commands \
     --lint-cmd "c: sh -c 'clang -fsyntax-only \"\$@\" > /tmp/aider_lint.log 2>&1; RC=\$?; cat /tmp/aider_lint.log | $SQZ_BIN; exit \$RC'" \
     --auto-test --test-cmd "sh -c '$TEST_RUNNER > /tmp/aider_test.log 2>&1; RC=\$?; cat /tmp/aider_test.log | $SQZ_BIN; cat /tmp/aider_test.log; exit \$RC'" \
     --git \
     --no-show-model-warnings --yes-always
fi


So, while the linuxlator is installed (see #30) it's quite nice to be able to run the latest version of claude.

Edit: 2026-09-30 Another way of running the latest version of claude

curl -fsSL https://claude.ai/install.sh -o /tmp/claude_install.sh


Edit the file /tmp/claude_install.sh and change the Detect platform to this or your preferred architecture:

sh:
# Detect platform
os="linux"
#case "$(uname -s)" in
#    Darwin) os="darwin" ;;
#    Linux) os="linux" ;;
#    MINGW*|MSYS*|CYGWIN*) echo "Windows is not supported by this script. See https://code.claude.com/docs for installation options." >&2; exit 1 ;;
#    *) echo "Unsupported operating system: $(uname -s). See https://code.claude.com/docs for supported platforms." >&2; exit 1 ;;
#esac
 
arch="x64"
#case "$(uname -m)" in
#    x86_64|amd64) arch="x64" ;;
#    arm64|aarch64) arch="arm64" ;;
#    *) echo "Unsupported architecture: $(uname -m)" >&2; exit 1 ;;
#esac

Run the install script
bash /tmp/install_claude.sh

You can add stable or latest or version number to above.

And of course these 2 steps
sudo mkdir -p /lib/x86_64-linux-gnu
sudo ln -sf /compat/linux/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 /lib/x86_64-linux-gnu/ld-linux-x86-64.so.2


It will install the latest version to $HOME/.local/bin


Edit: 2026-09-27 Alternate way to run latest claude or whatever version you want


cd $HOME
sudo npm install -g --allow-scripts=@anthropic-ai/claude-code @anthropic-ai/claude-code@2.1.281
cd /tmp
sudo npm pack @anthropic-ai/claude-code-linux-x64@2.1.281
sudo tar -xf anthropic-ai-claude-code-linux-x64-2.1.281.tgz
sudo mv package/claude /usr/local/lib/node_modules/@anthropic-ai/claude-code/bin/claude-linux-x64
sudo chmod +x /usr/local/lib/node_modules/@anthropic-ai/claude-code/bin/claude-linux-x64
sudo rm -rf package anthropic-ai-claude-code-linux-x64-2.1.281.tgz
sudo mkdir -p /lib/x86_64-linux-gnu
sudo ln -sf /compat/linux/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 /lib/x86_64-linux-gnu/ld-linux-x86-64.so.2
cd $HOME
ln -s /usr/local/lib/node_modules/@anthropic-ai/claude-code/bin/claude-linux-x64 /home/insertyourunsernamehere/.local/bin/claude
claude --version


Original method

This can be done by doing 3 steps providing you have node installed.
  1. sudo npm install -g --allow-scripts=@anthropic-ai/claude-code @anthropic-ai/claude-code
  2. sudo mkdir -p /lib/x86_64-linux-gnu
  3. sudo ln -sf /compat/linux/lib/x86_64-linux-gnu/ld-linux-x86-64.so.2 /lib/x86_64-linux-gnu/ld-linux-x86-64.so.2
What this does is that node grabs the latest version of claude and installs it to user profile. That version of claude has a dependency to a linux library. Since we have the linuxlator we have that library. Create the folder and symlink to trick claude into running.




/grandpa
 
Last edited:
sh:
#!/bin/sh

# ==============================================================================
# 🎛️ DSET CHAT TEMPLATE
# ==============================================================================

MODEL_LOWER=$(echo "$SELECTED_MODEL" | tr '[:upper:]' '[:lower:]')

# Global fallback sampling defaults
AIDER_ARCH="openai"
JINJA_FLAG="--jinja"
TEMPLATE_FLAG=""
LLAMA_PARALLEL="--parallel 1"
LLAMA_TOOLS=""

# ------------------------------------------------------------------------------
# MAPPING PER MODEL FAMILY
# ------------------------------------------------------------------------------

# High-precision deterministic baseline sampling optimized for code generation
SAMPLING_FLAGS="-n -1 --repeat-penalty 1.08 --repeat-last-n 512 --temp 0.3 --top-k 40 --top-p 0.90"

if echo "$MODEL_LOWER" | grep -iq "llama-3" || echo "$MODEL_LOWER" | grep -iq "llama3"; then
    STOP_FLAGS="-r <|eot_id|>,<|eom_id|>,<|im_end|>"
    AIDER_ARCH="llama3"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.10 --repeat-last-n 512 --temp 0.2 --top-k 40 --top-p 0.95"

elif echo "$MODEL_LOWER" | grep -iq "llama" || echo "$MODEL_LOWER" | grep -iq "codellama"; then
    TEMPLATE_FLAG="--chat-template llama2"
    STOP_FLAGS="-r </s>"
    AIDER_ARCH="llama2"
    JINJA_FLAG=""
    SAMPLING_FLAGS="-n 4096 --repeat-penalty 1.10 --repeat-last-n 256 --temp 0.3 --top-k 40 --top-p 0.90"

elif echo "$MODEL_LOWER" | grep -iq "qwen2.5-coder-14b"; then
    STOP_FLAGS="-r <|im_end|>, <|im_start|>,<|endoftext|>"
    AIDER_ARCH="chatml"
    AIDER_ALIAS_MODEL="openai/qwen2.5-coder:14b-instruct"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.08 --repeat-last-n 512 --temp 0.3 --top-k 40 --top-p 0.90"

elif echo "$MODEL_LOWER" | grep -iq "qwen3.8-27b"; then
    STOP_FLAGS="-r <|im_end|>,<|im_start|>,<|endoftext|>"
    AIDER_ARCH="chatml"
    SAMPLING_FLAGS="-n 16384 --repeat-penalty 1.05 --repeat-last-n 512 --temp 0.2 --top-k 5 --top-p 0.80"

elif echo "$MODEL_LOWER" | grep -iq "gemma4" || echo "$MODEL_LOWER" | grep -iq "gemma"; then
    if echo "$MODEL_LOWER" | grep -iq "gemma4"; then JINJA_FLAG="--jinja"; else TEMPLATE_FLAG="--chat-template gemma"; JINJA_FLAG=""; fi
    STOP_FLAGS="-r <|im_end|>,<end_of_turn>,<|endoftext|>"
    AIDER_ARCH="gemma"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.10 --repeat-last-n 512 --temp 0.3 --top-k 40 --top-p 0.90"

elif echo "$MODEL_LOWER" | grep -iq "granite-4.2-3b"; then
    STOP_FLAGS="-r <|im_end|>,<|end_of_turn|>,</think>,</s>"
    AIDER_ARCH="chatml"
    AIDER_ALIAS_MODEL="openai/qwen2.5-coder:7b-instruct"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.10 --repeat-last-n 512 --temp 0.1 --top-k 40 --top-p 0.90"

elif echo "$MODEL_LOWER" | grep -iq "granite-4.2-8b"; then
    STOP_FLAGS="-r <|im_end|>,<|end_of_turn|>,</think>,</s>"
    AIDER_ARCH="chatml"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.10 --repeat-last-n 512 --temp 0.3 --top-k 40 --top-p 0.90"

elif echo "$MODEL_LOWER" | grep -iq "mistral" || echo "$MODEL_LOWER" | grep -iq "deepseek"; then
    STOP_FLAGS="-r </s>,<|im_end|>"
    [ "$AIDER_ARCH" = "openai" ] && AIDER_ARCH="mistral"
    SAMPLING_FLAGS="-n 8192 --repeat-penalty 1.05 --repeat-last-n 256 --temp 0.2 --top-k 5 --top-p 0.80"

else
    STOP_FLAGS="-r <|im_end|>,</s>"
    AIDER_ARCH="chatml"
    SAMPLING_FLAGS="-n -1 --repeat-penalty 1.10 --repeat-last-n 512 --temp 0.3 --top-k 40 --top-p 0.90"
fi

# ------------------------------------------------------------------------------
# 3. AIDER PERSONA SPECIFICATION
# ------------------------------------------------------------------------------
if [ "$AIDER_ARCH" = "chatml" ] || [ "$AIDER_ARCH" = "llama3" ]; then
    AIDER_BODY="You are a senior FreeBSD C developer. Output the requested code file directly inside a standard markdown code block. Do not explain, talk, or philosophise."
else
    AIDER_BODY="You are an expert programmer. Always return your edits in the requested format immediately without preamble."
fi

if [ "$client_choice" = "2" ]; then
    STOP_FLAGS="-r </tool_call>,<|end|>,<|im_end|>,<end_of_turn>,</turn>,</s>"
fi

# Export parameters back to the background service shell daemon
export TEMPLATE_FLAG
export STOP_FLAGS
export LLAMA_PARALLEL
export LLAMA_TOOLS
export JINJA_FLAG
export SAMPLING_FLAGS
export AIDER_MODEL_ARCHITECTURE="$AIDER_ARCH"
export AIDER_BODY_PROMPT="$AIDER_BODY"
export AIDER_ALIAS_MODEL

if [ "${CONFIG_DEBUG:-false}" = "true" ]; then
    echo "🐛 [DEBUG set-chat_template.sh] Unified Configuration Loaded Contextually." >&2
fi
sh:
#!/bin/sh

# ==============================================================================
# 🚀 START LLAMA SERVER WITH CALCULATED CONTEXT AND HANDOVER TO AIDER/CLAUDE
# ==============================================================================
set -e

# Input from run-ai.sh
TARGET_MODEL_PATH="$1"
NGL="$2"
CACHE_TYPE="$3"
OFFLOAD_KQV="$4"
SERVER_HOST="$5"
SERVER_PORT="$6"
FINAL_CTX="$7"

if [ -z "$TARGET_MODEL_PATH" ] || [ -z "$NGL" ] || [ -z "$CACHE_TYPE" ] || [ -z "$FINAL_CTX" ]; then
    echo "❌ System Error: Missing critical deployment arguments." >&2
    exit 1
fi

TARGET_MODEL_NAME=$(basename "$TARGET_MODEL_PATH")
CONFIG_FILE="$HOME/.git-analyzer.conf"

# Baseline fallback parameter defaults
SCRIPTS_DIR="$HOME/scripts"
LLAMA_SERVER_BIN="/compat/linux/opt/llama-build/build/bin/llama-server"
DUMMY_UVM_HOOK="/compat/linux/opt/llama/bin/dummy-uvm.so"
SYSTEM_LOG_FILE="/var/log/llama-server.log"
DEBUG_MODE="false"

# Load global configuration directives dynamically if present
if [ -f "$CONFIG_FILE" ]; then
    . "$CONFIG_FILE"
    SCRIPTS_DIR="${CONFIG_SCRIPTS_PATH:-$HOME/scripts}"
    LLAMA_SERVER_BIN="${CONFIG_LLAMA_SERVER_BIN:-/compat/linux/opt/llama-build/build/bin/llama-server}"
    DUMMY_UVM_HOOK="${CONFIG_DUMMY_UVM_HOOK:-/compat/linux/opt/llama/bin/dummy-uvm.so}"
    DEBUG_MODE="${CONFIG_DEBUG:-false}"
fi

# ------------------------------------------------------------------------------
# 🤖 STEP 3: SELECT TARGET AI APPLICATION
# ------------------------------------------------------------------------------
echo ""
echo "=============================================================================="
echo " 🤖 STEP 3: SELECT TARGET AI APPLICATION"
echo "=============================================================================="
echo "  1) Aider"
echo "  2) Claude Code"
echo "=============================================================================="
printf " Choose frontend application [1-2]: "
read -r CLIENT_CHOICE
echo "========================================================================"

if [ "$CLIENT_CHOICE" != "1" ] && [ "$CLIENT_CHOICE" != "2" ]; then
    echo "❌ Invalid application choice. Aborting." >&2
    exit 1
fi

# ------------------------------------------------------------------------------
# ⏳ Step 4: Dynamically Compute Exact Safe -b (Batch) and -ub (Ubatch) Flags
# ------------------------------------------------------------------------------
echo ""
echo " ⏳ Step 4: Dynamically computing precise execution batch matrix..."

# Standardize the logical batch cap to match modern client expectations
CALCULATED_B=2048

# Parse structural model layers and attention metrics from the GGUF header dictionary
MODEL_STATS=$("$SCRIPTS_DIR/get-model_gguf_data.py" "$TARGET_MODEL_PATH" 2>/dev/null)
TOTAL_LAYERS=$(echo "$MODEL_STATS" | grep "Number of Layers:" | awk '{print $NF}')
HEAD_DIM=$(echo "$MODEL_STATS" | grep "Head Dimension:" | awk '{print $NF}')
KV_HEADS=$(echo "$MODEL_STATS" | grep "Head Count KV:" | awk '{print $NF}')
TRAINED_CTX_RAW=$(echo "$MODEL_STATS" | grep "Trained Context:" | awk '{print $NF}')
EMBEDDING_LEN=$(echo "$MODEL_STATS" | grep "Embedding Length:" | awk '{print $NF}' | tr -d '\n\r ')

[ -z "$EMBEDDING_LEN" ] || [ "$EMBEDDING_LEN" = "Unknown" ] && EMBEDDING_LEN=4096

# Get VRAM and RAM from script
HW_OUTPUT=$("$SCRIPTS_DIR/check-vram_ram.sh" 2>/dev/null)
TOTAL_VRAM_BYTES=$(echo "$HW_OUTPUT" | grep "Dedicated VRAM" | awk '{print $NF}' | sed 's/[^0-9]//g')
[ -z "$TOTAL_VRAM_BYTES" ] && TOTAL_VRAM_BYTES=12884901888

# Gather model file footprint metrics
FILE_SIZE_BYTES=$(stat -f %z "$TARGET_MODEL_PATH" 2>/dev/null || stat -c %s "$TARGET_MODEL_PATH")

# Compute the exact physical size of the chosen KV Cache allocation (Bytes)
case "$CACHE_TYPE" in
    f16)  bytes_per_elem=2.0 ;;
    q8_0) bytes_per_elem=1.0625 ;;
    q4_0) bytes_per_elem=0.5625 ;;
    *)    bytes_per_elem=0.5625 ;;
esac
KV_CACHE_BYTES=$(awk -v f="$FINAL_CTX" -v l="${TOTAL_LAYERS:-40}" -v kh="${KV_HEADS:-32}" -v hd="${HEAD_DIM:-128}" -v bpe="$bytes_per_elem" 'BEGIN {print 2 * f * l * kh * hd * bpe}')

# Isolate the exact net remaining VRAM space left for the compute graph
USER_OVERHEAD_BYTES=$(awk -v oh="${CONFIG_VRAM_OVERHEAD:-1.0}" 'BEGIN {print oh * 1024 * 1024 * 1024}')
REMAINING_VRAM_FOR_GRAPH=$(awk -v tv="$TOTAL_VRAM_BYTES" -v oh="$USER_OVERHEAD_BYTES" -v ms="$FILE_SIZE_BYTES" -v kv="$KV_CACHE_BYTES" 'BEGIN {val = tv - oh - ms - kv; print (val < 0 ? 0 : val)}')

# Reverse the tensor activation formula to calculate the absolute max safe ubatch token size
RAW_MAX_UB=$(awk -v free="$REMAINING_VRAM_FOR_GRAPH" -v h="$EMBEDDING_LEN" -v l="${TOTAL_LAYERS:-40}" 'BEGIN {denom = h * 16 * l; print (denom > 0 ? int(free / denom) : 0)}')

# Align the raw number down to the nearest standard stable binary power-of-two block
# Keep minimum ub at 128 to avoid too slow performance
if [ "$OFFLOAD_KQV" = "false" ] || [ "$RAW_MAX_UB" -lt 64 ]; then
    CALCULATED_UB=64
elif [ "$RAW_MAX_UB" -ge 1024 ]; then
    CALCULATED_UB=1024
elif [ "$RAW_MAX_UB" -ge 512 ]; then
    CALCULATED_UB=512
elif [ "$RAW_MAX_UB" -ge 256 ]; then
    CALCULATED_UB=256
else
    CALCULATED_UB=128
fi

echo "📊 Dynamic Batch Locked : Logical Batch (-b): $CALCULATED_B | Physical Micro-Batch (-ub): $CALCULATED_UB"

# Handle standard fallback limits for sequence thresholds
if [ -z "$TRAINED_CTX_RAW" ] || [ "$TRAINED_CTX_RAW" = "Unknown" ] || [ "$TRAINED_CTX_RAW" -le 0 ]; then
    TRAINED_CTX=32768
else
    TRAINED_CTX="$TRAINED_CTX_RAW"
fi

ROPE_FLAGS=""
if [ "$FINAL_CTX" -gt "$TRAINED_CTX" ] 2>/dev/null; then
    SCALE_FACTOR=$(awk -v f="$FINAL_CTX" -v t="$TRAINED_CTX" 'BEGIN {print f / t}')
    FREQ_SCALE=$(awk -v sf="$SCALE_FACTOR" 'BEGIN {print 1 / sf}')
    ROPE_FLAGS="--rope-scaling yarn --rope-freq-scale $FREQ_SCALE"
    echo "📊 RoPE Extension Active : Context stretched by ${SCALE_FACTOR}x ($TRAINED_CTX --> $FINAL_CTX)"
else
    echo "📊 RoPE Extension Active : No (Context fits perfectly within model boundaries)"
fi
export ROPE_FLAGS

# ------------------------------------------------------------------------------
# Step 5: Get chat template values
# ------------------------------------------------------------------------------
echo "🚀 Calling set-chat_template.sh..."
export SELECTED_MODEL="$TARGET_MODEL_NAME"
export client_choice="$CLIENT_CHOICE"

if [ -f "$SCRIPTS_DIR/set-chat_template.sh" ]; then
    . "$SCRIPTS_DIR/set-chat_template.sh"
else
    echo "❌ System Error: $SCRIPTS_DIR/set-chat_template.sh not found!" >&2
    exit 1
fi
# ------------------------------------------------------------------------------
# Step 6: Log profiles and launch llama-server backend daemon
# ------------------------------------------------------------------------------

if [ "$DEBUG_MODE" = "true" ]; then
    echo ""
    echo "=============================================================================="
    echo " 🔬 DEBUG: INCOMING MATRIX ARGUMENTS VERIFICATION"
    echo "=============================================================================="
    echo "  • Passed Model Path   : $TARGET_MODEL_PATH"
    echo "  • Computed GPU Layers : $NGL"
    echo "  • Selected Cache Type : $CACHE_TYPE"
    echo "  • Target Network Host : $SERVER_HOST"
    echo "  • Target Network Port : $SERVER_PORT"
    echo "  • Extracted Max Ctx   : $FINAL_CTX tokens"
    echo "  • Allocated Batch / UB: $CALCULATED_B / $CALCULATED_UB"
    echo "  • Est. KV Cache Space : $(awk -v b="$KV_CACHE_BYTES" 'BEGIN {printf "%.2f MB", b/1024/1024}')"
    echo "  • Remaining Graph Room: $(awk -v b="$REMAINING_VRAM_FOR_GRAPH" 'BEGIN {printf "%.2f MB", b/1024/1024}')"
    echo "=============================================================================="
    echo ""
fi

echo "🚀 Initializing llama-server backend in the background..."

if [ "$OFFLOAD_KQV" = "false" ]; then
    KQV_FLAG="--no-kv-offload"
else
    KQV_FLAG=""
fi

nohup env \
  PATH="/compat/linux/usr/local/cuda-12.8/bin:$PATH" \
  CUDAToolkit_ROOT="/compat/linux/usr/local/cuda-12.8" \
  LD_LIBRARY_PATH="/compat/linux/usr/local/cuda-12.8/lib64:/usr/lib64:/usr/lib/x86_64-linux-gnu" \
  LD_PRELOAD="$DUMMY_UVM_HOOK" \
  "$LLAMA_SERVER_BIN" \
    --ctx-size "$FINAL_CTX" \
    --n-gpu-layers "$NGL" \
    -b "$CALCULATED_B" \
    -ub "$CALCULATED_UB" \
    $KQV_FLAG \
    $ROPE_FLAGS \
    -fa on \
    -v \
    --cont-batching \
    --slots \
    -ctk "$CACHE_TYPE" \
    -ctv "$CACHE_TYPE" \
    $SAMPLING_FLAGS \
    $JINJA_FLAG \
    $TEMPLATE_FLAG \
    $LLAMA_PARALLEL \
    --host "$SERVER_HOST" \
    --port "$SERVER_PORT" \
    $LLAMA_TOOLS \
    -m "$TARGET_MODEL_PATH" > "$SYSTEM_LOG_FILE" 2>&1 &

SERVER_PID=$!

# Save the active background process ID securely to a PID file for global orchestration tracking
echo "$SERVER_PID" > /tmp/llama-server.pid

# ------------------------------------------------------------------------------
# 🤝 Step 7: Handshake Verification Loop
# ------------------------------------------------------------------------------
echo "⏳ Waiting for llama-server handshake at http://$SERVER_HOST:$SERVER_PORT/health ..."

PROBE_HOST="$SERVER_HOST"
if [ "$PROBE_HOST" = "0.0.0.0" ]; then
    PROBE_HOST="127.0.0.1"
fi

max_attempts=360
attempt=1
success=0

while [ "$attempt" -le "$max_attempts" ]; do
    if ! kill -0 "$SERVER_PID" 2>/dev/null; then
        echo "❌ Error: llama-server process terminated unexpectedly. Inspect log trace at: $SYSTEM_LOG_FILE" >&2
        exit 1
    fi

    if fetch -q -T 3 -o - "http://$PROBE_HOST:$SERVER_PORT/health" 2>/dev/null | grep -q '"status":'; then
        success=1
        break
    fi

    printf "."
    sleep 2
    attempt=$((attempt + 1))
done

echo ""
if [ "$success" -eq 1 ]; then
    echo "✅ Handshake successful! llama-server is online and fully operational."
    echo "🚀 Environment pipeline deployment complete."
else
    echo "❌ Error: Timeout waiting for llama-server handshake. Check state log: $SYSTEM_LOG_FILE" >&2
    exit 1
fi

# ------------------------------------------------------------------------------
# Step 8: Handoff Execution to Selected Frontend Client
# ------------------------------------------------------------------------------
echo ""
if [ "$CLIENT_CHOICE" = "1" ]; then
    echo "⚡ Handoff pipeline execution directly to Aider..."
 
    AIDER_EXEC_RAW="${CONFIG_AIDER_PATH:-aider}"
    CLIENT_HOST="$SERVER_HOST"
    if [ "$CLIENT_HOST" = "0.0.0.0" ]; then
        CLIENT_HOST="127.0.0.1"
    fi
 
    AIDER_EXEC_WITH_BASE="$AIDER_EXEC_RAW --openai-api-base http://$CLIENT_HOST:${SERVER_PORT}/v1"
 
    exec "$SCRIPTS_DIR/run-aider.sh" \
        "$TARGET_MODEL_NAME" \
        "$FINAL_CTX" \
        "$AIDER_EXEC_WITH_BASE" \
        "$FINAL_CTX"
    
elif [ "$CLIENT_CHOICE" = "2" ]; then
    echo "⚡ Handoff pipeline execution directly to Claude Code..."
 
    CLAUDE_EXEC_PATH="${CONFIG_CLAUDE_PATH:-claude}"
 
    exec "$SCRIPTS_DIR/run-claude.sh" \
        "$TARGET_MODEL_NAME" \
        "$FINAL_CTX" \
        "$CLAUDE_EXEC_PATH" \
        "$FINAL_CTX"
fi
sh:
#!/bin/sh
# test_runner.sh

# 1. Tvinga Qt/KDE att hålla tyst under körningen för att spara Aiders tokens
export QT_LOGGING_RULES="*.debug=false;*.warning=false;kf.*.warning=false"

# 2. Hitta ALLA .c-filer i nuvarande mapp samt i src/ underkatalogen
ALL_FILES=$(find . -maxdepth 2 -name "*.c" -o -path "./src/*.c" 2>/dev/null | tr '\n' ' ')

if [ -n "$ALL_FILES" ]; then
    # 3. Välj ett namn för binären baserat på den senast ändrade C-filen i git
    MAIN_FILE=$(git diff-tree --no-commit-id --name-only -r HEAD | grep "\.c$" | head -n 1)
    if [ -z "$MAIN_FILE" ] || [ ! -f "$MAIN_FILE" ]; then
        # Fallback: Ta den fysiskt senast modifierade C-filen på disk (bland alla hittade)
        MAIN_FILE=$(ls -t $ALL_FILES 2>/dev/null | head -n 1)
    fi
 
    # Extrahera bara filnamnet utan sökväg och ändelse för binären
    BASE_NAME=$(basename "$MAIN_FILE")
    BINARY="${BASE_NAME%.c}"
    if [ -z "$BINARY" ]; then BINARY="ai_test_build"; fi
 
    echo "⚙️ Automatically compiling all discovered files ($ALL_FILES) into binary [$BINARY] using clang..."
 
    # 4. Kompilera alla filer tillsammans och kör binären (släng dolda system-felmeddelanden)
    clang $ALL_FILES -o "$BINARY" && ./"$BINARY" 2>/dev/null
else
    echo "❌ Error: Aider test runner couldn't locate any active C files in . or ./src."
    exit 1
fi

/grandpa
 
Last edited:
Great topic for a thread, thanks cracauer@.

Here's some technical aspects for running AI models locally on unbalanced crap hardware with llama-cpp, some random AI models from huggingface and aider.

The crap Hardware:

CPU Intel(R) Core(TM) i5-2500 CPU @ 3.30GHz ~ 2011​
16 Gb DD3 RAM @ 1600 ~ 2007​
RTX 3060 with 12 Gb NVRAM ~ 2021​

Installations:

llama-cpp
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
mkdir build
cd build
cmake -DGGML_VULKAN=1 ..
cmake --build . --config Release -j 3
cp llama* lib*.so* ~/.local/bin/


aider - this one is a bit trickier - git clone the source and then pip install in a venv and fix the dependency errors - it can be done but requires some kerfuffling​

Task for AI to do - the prompt, the testcase, whatever you want to call it. The goal is to make AI write the code and test until all errors are gone without intervention.

Create a new C file with a descriptive name. Inside it, implement a function to read the value for the CPU temperature using sysctlbyname with the modern path "dev.cpu.0.temperature". Note that FreeBSD returns this value as an integer in deci-Kelvin (e.g., 3000 means 300.0 Kelvin). Convert this value to Celsius by dividing by 10.0 and subtracting 273.15, then print it out in main() Also in the same code read the values for load average and print those out. If the compilation fails, intercept the errors and fix them.​

The random models and how they did.

1)4.36Gqwen2.5-coder-7b-instruct-q4_k_m.ggufsuccess for load average but fail for temperature~ 54 t/s
2)6.87Ggemma-4-12b-Q4_K_M.gguffails caught in a reasoning loop~ 29 t/s
3)6.87Ggemma4-coding-Q4_K_M.ggufsuccess~ 30 t/s
4)7.95GLlama-3-Hercules-5.1-8B-Q8_0.gguffails to create the file~ 35 t/s
5)8.01GCodestral-22B-v0.1-IQ3_XXS.ggufsuccess~ 18 t/s
6)8.37Gqwen2.5-coder-14b-instruct-q4_k_m.ggufsuccess for temperature fail for load averages~ 30 t/s
7)8.43Gmicrosoft_Phi-4-reasoning-Q4_K_M.gguffails by hanging when fixing error~ 32 t/s
8)8.65GDevstral-Small-2-24B-Instruct-2512-UD-Q2_K_XL.ggufsuccess for temperature fails for load average~ 19 t/s
9)8.74GNorth-Mini-Code-1.0-UD-IQ1_M.ggufsuccess~ 54 t/s
a)8.91GQwen3.6-27B-UD-IQ2_XXS.gguffails~ 18 t/s
b)9.11Ggemma4-coding-Q6_K.ggufsuccess for temperature fails for load averages~ 25 t/s
c)9.11GQwen3.5-9B-Q8_0.ggufsuccess for temperature fails for load averages~ 25 t/s
d)14.84GQwen2.5-Coder-32B-Instruct-Q3_K_M.ggufsuccess for temperature fails slightly for load averages (2/3)~ 1 t/s

Warning:
If you run the below scripts you are responsible. You will see llama-cpp warnings like below so make sure you are safe.
srv llama_server: -----------------
srv llama_server: CORS is set to allow all origins ('*') and no API key is set
srv llama_server: this can be a security risk (cross-origin attacks)
srv llama_server: more info: https://github.com/ggml-org/llama.cpp/pull/25655
srv llama_server: -----------------
srv llama_server: -----------------
srv llama_server: the following feature(s) are enabled:
srv llama_server: server tools (experimental)
srv llama_server: do not expose the server to untrusted environments
srv llama_server: -----------------


Comments:
- llama-cpp was installed from ports first, that might have helped to build it from source
- for aider the dependencies that caused errors were installed as pkg and then edited out from requirements.txt
- the reason speed drops for models sized > NVRAM is that the CPU/RAM slows everything down when offloading from NVRAM is being done
- there is margin in the memory management in llama startup script to allow kde and chromium to run at the same time as the ai model
- there is a feeling that Vulkan not always is giving memory back after stopping llama-server with ctrl-c

Edit:
The script below has been tweaked and updated.

This script will run models up to ~19GB in size for any crap old hardware with 16GB RAM and 12GB NVRAM which means 32/34B models with Q4_K_M/Q3_K_M. The speed is estimated to around 0.5-1.5 t/s for any model that goes above 12GB. For smaller models speed is estimated to around 20-50 t/s.

Script to start llama and pick a model:

sh:
#!/bin/sh

MODELPATH="$HOME/AI-models"
TOTAL_VRAM=12.0      # Physical VRAM in GB
MAX_LOCKED_RAM=10.0  # FreeBSD limit for mlock in GB

# Path to your working Python 3.12 binary inside your venv
PY_BIN="$HOME/venvs/aider-env/bin/python3"

run_llama() {
    FILE_NAME="$1"
    FILE_SIZE_GB="$2"
    FULL_PATH="$MODELPATH/$FILE_NAME"

    echo "──────────────────────────────────────────────────────"
    echo "📊 Analyzing $FILE_NAME ($FILE_SIZE_GB GB)..."

    # Verify NVIDIA GPU and driver status cleanly
    echo "🎮 Verifying NVIDIA GPU status..."
    if [ ! -c /dev/nvidia0 ]; then
        echo "❌ Error: NVIDIA device node (/dev/nvidia0) is missing!"
        return 1
    fi

    NV_VERSION=$(sysctl -n hw.nvidia.version 2>/dev/null)
    if [ -z "$NV_VERSION" ]; then
        echo "❌ Error: NVIDIA kernel module is not active!"
        return 1
    fi
    echo "✅ GPU Status: NVIDIA Graphics Driver ($NV_VERSION) is active."

    echo "🧠 Reading model built-in metadata live..."

    export TARGET_GGUF_PATH="$FULL_PATH"

    METADATA=$("$PY_BIN" -c '
import sys, struct, os

def read_gguf_metadata(filepath):
    ctx_val, blocks_val = None, None
    try:
        with open(filepath, "rb") as f:
            # Verify Magic Header "GGUF"
            magic = f.read(4)
            if magic != b"GGUF": return None, None
        
            # Read Version (uint32) and V2/V3 fields
            version = struct.unpack("<I", f.read(4))[0]
            if version not in [2, 3]: return None, None
        
            # Read counts
            tensor_count = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
            kv_count = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
        
            # Types map according to GGUF spec
            # 0=uint8, 1=int8, 2=uint16, 3=int16, 4=uint32, 5=int32, 6=float32, 7=bool, 8=string, 9=array...
            for _ in range(kv_count):
                # Read key string
                key_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
                key = f.read(key_len).decode("utf-8", errors="ignore")
            
                # Read value type
                val_type = struct.unpack("<I", f.read(4))[0]
            
                # Helper to skip or read values based on type
                def skip_value(t):
                    if t in [0, 1, 7]: f.seek(1, 1)
                    elif t in [2, 3]: f.seek(2, 1)
                    elif t in [4, 5, 6]: return struct.unpack("<I", f.read(4))[0]
                    elif t == 8: # String
                        s_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
                        f.seek(s_len, 1)
                    elif t == 9: # Array
                        arr_type = struct.unpack("<I", f.read(4))[0]
                        arr_len = struct.unpack("<Q" if version == 3 else "<I", f.read(8 if version == 3 else 4))[0]
                        for _ in range(arr_len): skip_value(arr_type)
                    elif t in [10, 11, 12, 13]: f.seek(8, 1) # uint64, int64, float64
                    return None

                res = skip_value(val_type)
            
                # Capture our specific targets cleanly
                if "context_length" in key and res:
                    ctx_val = res
                elif "block_count" in key and res:
                    blocks_val = res
                
                if ctx_val and blocks_val:
                    break
    except Exception:
        pass
    return ctx_val, blocks_val

ctx, blocks = read_gguf_metadata(os.environ.get("TARGET_GGUF_PATH", ""))
ctx_str = str(ctx) if ctx else ""
blocks_str = str(blocks) if blocks else ""
print(ctx_str + " " + blocks_str)
')

    TRAINED_CTX=$(echo "$METADATA" | awk '{print $1}')
    ACTUAL_LAYERS=$(echo "$METADATA" | awk '{print $2}')

    if [ -z "$TRAINED_CTX" ] || [ "$TRAINED_CTX" -lt 2048 ]; then
        TRAINED_CTX=32768
        echo "ℹ  Could not read trained context length. Using safe fallback: $TRAINED_CTX"
    else
        echo "🧠 Model is natively trained for a max of: $TRAINED_CTX tokens."
    fi

    if [ -z "$ACTUAL_LAYERS" ] || [ "$ACTUAL_LAYERS" -lt 10 ]; then
        ACTUAL_LAYERS=64
        echo "ℹ  Could not read layer count. Using baseline fallback: $ACTUAL_LAYERS layers."
    else
        echo "🧱 Model has an architectural structure of: $ACTUAL_LAYERS layers."
    fi

#
# MEMORY MANAGEMENT (DYNAMIC VRAM/RAM SPLIT & DYNAMIC KV CACHE)
#
    EVAL_OPTS=$(awk -v vram="$TOTAL_VRAM" -v ram="$MAX_LOCKED_RAM" -v size="$FILE_SIZE_GB" -v max_ctx="$TRAINED_CTX" -v total_layers="$ACTUAL_LAYERS" '
    BEGIN {
        system_vram_overhead = 1.8;
        available_vram = vram - system_vram_overhead;
    
        # SCENARIO 1: Model fits entirely in VRAM -> Use full quality f16 KV Cache
        if (size <= available_vram) {
            cache_type = "f16";
            cache_factor = 0.0065; # Full precision footprint
        
            vram_left_for_cache = available_vram - size;
            calculated_ctx = int((vram_left_for_cache * 1000) / (total_layers * cache_factor));
            ctx = int(calculated_ctx / 1024) * 1024;
        
            if (ctx < 2048) ctx = 2048;
            if (ctx > max_ctx) ctx = max_ctx;
        
            ngl = total_layers + 5;
            ram_fallback = 0;
            b = 512; ub = 256;
            mode = "mlock"; # Safe to lock in RAM since 0 model layers spill over
        }
        # SCENARIO 2: Model requires split-mode -> Switch to q4_0 KV Cache to save VRAM for layers
        else {
            cache_type = "q4_0";
            cache_factor = 0.0017; # Compressed footprint
        
            ctx = 4096;
            if (ctx > max_ctx) ctx = max_ctx;
        
            cache_vram_cost = (ctx / 1000) * total_layers * cache_factor;
            usable_vram_for_layers = available_vram - cache_vram_cost;
        
            if (usable_vram_for_layers < 0) usable_vram_for_layers = 0;
        
            vram_ratio = usable_vram_for_layers / size;
            ngl = int(total_layers * vram_ratio);
        
            if (ngl < 0) ngl = 0;
            if (ngl > (total_layers + 5)) ngl = total_layers + 5;
        
            ram_fallback = size - usable_vram_for_layers;
            b = 512; ub = 256;
            mode = "mmap"; # Use standard mmap to safely spill past FreeBSD mlock limits
        }
    
        if (ram_fallback > ram) {
            print "ERROR: Model requires " ram_fallback " GB system RAM offload, which exceeds your 10 GB limit!" > "/dev/stderr";
            printf "0 ERROR 0 0 0 error\n";
            exit 1;
        }
    
        printf "%d %s %d %d %d %s\n", ctx, cache_type, b, ub, ngl, mode;
    }')

    read -r ctx type b ub ngl mode <<EOF
$EVAL_OPTS
EOF

    if [ "$type" = "ERROR" ] || [ -z "$ctx" ] || [ "$ctx" -eq 0 ]; then
        echo "❌ ERROR: Launch Aborted! Constraints not met."
        return 1
    fi

    # Calculate safe prompt margin
    SAFE_PROMPT_LIMIT=$((ctx - 512))
    if [ "$SAFE_PROMPT_LIMIT" -lt 1024 ]; then
        SAFE_PROMPT_LIMIT=$ctx
    fi

    echo "⚙️  Optimized profile: CTX=$ctx, NGL=$ngl, BATCH=$b, UBATCH=$ub, CACHE=$type, MODE=$mode"
    echo "🔄 Context Shifting: Active (Infinite scrolling chat enabled)"
    echo "💡 Safe Prompt Limit: Do not exceed $SAFE_PROMPT_LIMIT tokens in a SINGLE prompt."
    echo "──────────────────────────────────────────────────────"

    # Start the server with q4_0 KV cache flags and flash attention
    llama-server \
        -t 3 \
        -b "$b" \
        -tb 4 \
        -ub "$ub" \
        -c "$ctx" \
        --context-shift \
        -fa on \
        --load-mode "$mode" \
        -ctk "$type" \
        -ctv "$type" \
        --jinja \
        --host 0.0.0.0 \
        --port 8080 \
        --alias kAI \
        --temp 0.4 \
        --top-p 0.1 \
        --min-p 0.05 \
        --tools all \
        -m "$FULL_PATH"
}

# Check if model directory exists
if [ ! -d "$MODELPATH" ]; then
    echo "❌ Error: Model directory $MODELPATH does not exist."
    exit 1
fi

echo "Pick an AI model to run: "

index=1
chars="123456789abcdefghijklmnopqrstuvwxyz"

MATCHED_MODELS=$(find "$MODELPATH" -maxdepth 1 -name "*.gguf" 2>/dev/null | while read -r path; do
    bytes=$(stat -f %z "$path")
    name=$(basename "$path")
    echo "$bytes $name"
done | sort -n)

old_ifs=$IFS
IFS='
'
for line in $MATCHED_MODELS; do
    IFS=$old_ifs
    bytes=$(echo "$line" | awk '{print $1}')
    name=$(echo "$line" | cut -d' ' -f2-)
    size_gb=$(awk -v b="$bytes" 'BEGIN {printf "%.2f", b / 1024 / 1024 / 1024}')
    selector=$(echo "$chars" | cut -c "$index")
    if [ -z "$selector" ]; then
        break
    fi
    eval "MENU_KEY_$selector=\"$name\""
    eval "MENU_SIZE_$selector=\"$size_gb\""
    printf "%s)   %6sG   %s\n" "$selector" "$size_gb" "$name"
    index=$((index + 1))
    IFS='
'
done
IFS=$old_ifs
echo "0) Exit"

printf "Your choice: "
read choice

if [ "$choice" = "0" ] || [ -z "$choice" ]; then
    echo "Exiting."
    exit 0
fi

eval "SELECTED_MODEL=\$MENU_KEY_$choice"
eval "SELECTED_SIZE=\$MENU_SIZE_$choice"

run_llama "$SELECTED_MODEL" "$SELECTED_SIZE"


/grandpa
FWIW, the lama_run.sh (my name) script fails if the models are stored in a directory tree with depth > 1.
Example:
Code:
tingo@locaal:~ $ find $HOME/.cache//huggingface/hub -name "*.gguf"
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/Qwen_Qwen3.5-27B-Q6_K_L.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/mmproj-Qwen_Qwen3.5-27B-bf16.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/Qwen_Qwen3.5-27B-Q3_K_L.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/Qwen_Qwen3.5-27B-Q3_K_M.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/Qwen_Qwen3.5-27B-Q3_K_S.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen_Qwen3.5-27B-GGUF/snapshots/d7b113c40283f4d99f4eb0ec20d126ad653cc736/Qwen_Qwen3.5-27B-Q2_K_L.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--TheDrummer_Orion-26B-A4B-v1.1-GGUF/snapshots/ecc1d958b6b4a83445ab80262965e151e00d97ac/TheDrummer_Orion-26B-A4B-v1.1-IQ3_XS.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--TheDrummer_Orion-26B-A4B-v1.1-GGUF/snapshots/ecc1d958b6b4a83445ab80262965e151e00d97ac/mmproj-TheDrummer_Orion-26B-A4B-v1.1-bf16.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen3.8-27B-GGUF/snapshots/125a02af4987b57c7deb88d7f2ec58a5725c07c0/Qwen3.8-27B-IQ2_XS.gguf
/home/tingo/.cache//huggingface/hub/models--bartowski--Qwen3.8-27B-GGUF/snapshots/125a02af4987b57c7deb88d7f2ec58a5725c07c0/mmproj-Qwen3.8-27B-bf16.gguf
script, modified to use my paths, and I have taken out the 'maxdepth 1' from the find command
when run
Code:
tingo@locaal:~ $ sh bin/run_llama.sh
Pick an AI model to run: 
1)     0.00G   Qwen3.8-27B-IQ2_XS.gguf
2)     0.00G   Qwen_Qwen3.5-27B-Q2_K_L.gguf
3)     0.00G   Qwen_Qwen3.5-27B-Q3_K_L.gguf
4)     0.00G   Qwen_Qwen3.5-27B-Q3_K_M.gguf
5)     0.00G   Qwen_Qwen3.5-27B-Q3_K_S.gguf
6)     0.00G   Qwen_Qwen3.5-27B-Q6_K_L.gguf
7)     0.00G   TheDrummer_Orion-26B-A4B-v1.1-IQ3_XS.gguf
8)     0.00G   mmproj-Qwen3.8-27B-bf16.gguf
9)     0.00G   mmproj-Qwen_Qwen3.5-27B-bf16.gguf
a)     0.00G   mmproj-TheDrummer_Orion-26B-A4B-v1.1-bf16.gguf
0) Exit
Your choice: 1
──────────────────────────────────────────────────────
📊 Analyzing Qwen3.8-27B-IQ2_XS.gguf (0.00 GB)...
🎮 Verifying NVIDIA GPU status...
✅ GPU Status: NVIDIA Graphics Driver (NVIDIA UNIX x86_64 Kernel Module  595.84  Wed Jun 10 21:13:57 UTC 2026) is active.
🧠 Reading model built-in metadata live...
ℹ  Could not read trained context length. Using safe fallback: 32768
ℹ  Could not read layer count. Using baseline fallback: 64 layers.
⚙️  Optimized profile: CTX=32768, NGL=69, BATCH=512, UBATCH=256, CACHE=f16, MODE=mlock
🔄 Context Shifting: Active (Infinite scrolling chat enabled)
💡 Safe Prompt Limit: Do not exceed 32256 tokens in a SINGLE prompt.
──────────────────────────────────────────────────────
0.00.159.059 I log_info: verbosity = 3 (adjust with the `-lv N` CLI arg)
0.00.159.072 I device_info:
0.00.159.143 I   - Vulkan0 : NVIDIA GeForce RTX 5060 Ti (16557 MiB, 16071 MiB free)
0.00.159.150 I   - CPU     : CPU (32659 MiB, 32659 MiB free)
0.00.159.209 I system_info: n_threads = 3 (n_threads_batch = 4) / 12 | CPU : OPENMP = 1 | REPACK = 1 | 
0.00.159.215 I srv  llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true
0.00.159.238 I srv          init: running without SSL
0.00.159.270 I srv          init: using 11 threads for HTTP server
0.00.159.388 W srv  llama_server: -----------------
0.00.159.391 W srv  llama_server: Built-in tools are enabled, do not expose server to untrusted environments
0.00.159.392 W srv  llama_server: This feature is EXPERIMENTAL and may be changed in the future
0.00.159.392 W srv  llama_server: -----------------
0.00.159.398 I srv         start: binding port with default address family
0.00.160.545 I srv  llama_server: loading model
0.00.160.552 I srv    load_model: loading model '/home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf'
0.00.160.629 I common_init_result: fitting params to device memory ...
0.00.160.631 I common_init_result: (for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on)
0.00.160.664 E gguf_init_from_file: failed to open GGUF file '/home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf' (No such file or directory)
0.00.160.727 E llama_model_load: error loading model: llama_model_loader: failed to load model from /home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf
0.00.160.737 E llama_model_load_from_file_impl: failed to load model
0.00.160.771 E common_fit_params: encountered an error while trying to fit params to free device memory: failed to load model
0.00.160.779 E gguf_init_from_file: failed to open GGUF file '/home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf' (No such file or directory)
0.00.160.800 E llama_model_load: error loading model: llama_model_loader: failed to load model from /home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf
0.00.160.803 E llama_model_load_from_file_impl: failed to load model
0.00.160.804 E common_init_from_params: failed to load model '/home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf'
0.00.160.813 E srv    load_model: failed to load model, '/home/tingo/.cache/huggingface/hub/Qwen3.8-27B-IQ2_XS.gguf'
0.00.160.814 I srv    operator(): operator(): cleaning up before exit...
0.00.161.110 E srv  llama_server: exiting due to model loading error
I'm too tired to fix it now. Maybe another night.
 
FWIW, the lama_run.sh (my name) script fails if the models are stored in a directory tree with depth > 1.

tingo@locaal:~ $ sh bin/run_llama.sh
Pick an AI model to run:
1) 0.00G Qwen3.8-27B-IQ2_XS.gguf
2) 0.00G Qwen_Qwen3.5-27B-Q2_K_L.gguf

ℹ Could not read trained context length. Using safe fallback: 32768
ℹ Could not read layer count. Using baseline fallback: 64 layers.

I'm too tired to fix it now. Maybe another night.

Yes, the scripts are particular to my environment and I see you are not getting the file sizes nor the gguf data.

The scripts have evolved to the ones in post #32, #33 where they are maybe just a little bit cleaner. Sorry for any inconvenience.

The later versions are more split up so the gguf reader is now a script of its own. If you want to run with Vulkan you will need to modify the command-to-run and the path to the llama-cpp executable in ai-launcher.sh.

In an attempt to simplify the scripts should be callable with parameters so you can test their output by calling them with actual values for the params.

/grandpa
 
So, while the linuxlator is installed (see #30) it's quite nice to be able to run the latest version of claude.

Talking to myself again.

Here's a technical aspect. It's very difficult to get claude to write to disk if you are low on NVRAM and RAM.

It seems claude wants to use - relatively speaking - large context windows and also there's a general feeling that claude veers towards the anthropic models.

Even with environment variables like --permission-mode bypassPermissions, as large context window as possible, running from a bash terminal and a .git subdirectory it's neigh on impossible to make claude create new files on local disk. It can - sometimes - compile source code and then create executables and it can read files, just not create new ones.

If anyone finds a solution to making claude create new files on systems with low RAM/NVRAM please let me know.

/grandpa
 
Here's a very crude and rough guide for how to pick an AI-model to run locally with purpose to get as much out of the model as possible.

mysterious quant in file namesomething about equivalent bitsnumber of parameters the model's been trained on (B)
BF16 / FP16161-8
Q8_0 / INT881-8
Q6_K / INT66.51-8
Q5_K_M5.51-8
Q5_K_S / Q5_051-14
GPTQ414-32
IQ4_NL4.514-32
Q4_K_M4.514-32
IQ4_XS4.2514-70
Q4_K_S / INT4432-70
IQ3_M3.570+
Q3_K_L 3.570+
IQ3_S / IQ3_XXS370-120
Q3_K_M3.370-120
IQ2 / Q2270+
IQ11.5100+

If you download a model with Q8_0 in the filename then look for a model that has 1B up to 8B in the filename. Or if you download a model with IQ2 then look for 70B and upwards.

/grandpa
 
Another aspect for stubborn old fools like me trying to run AI models on hardware that has parts from 2011 ...

Here' s a neat tool to catch-and-compress the context window before it's being sent to the model.

In effect you can use a context that is up to 3 times larger than the one llama-server has allocated in NVRAM or RAM without llama knowing.

This is the tool on github - sqz

Installation and setup (you will to have rust installed):
cargo install sqz-cli
sqz init --global


Claude should pick up this by automagic but for aider the script run-aider.sh in #33 needs to be changed:
sh:
# ==============================================================================
# 2. APPLY HARD OPTIMISATION AND MEMORY RESTRICTIONS (SQZ ADAPTIVE CHECK)
# ==============================================================================
export AIDER_MAP_TOKENS="0"
export AIDER_MAX_CHAT_HISTORY_TOKENS="8192"
export AIDER_TEMPERATURE="0.1"

# Check if sqz is installed
if [ -x "$HOME/.cargo/bin/sqz" ]; then
    SQZ_BIN="$HOME/.cargo/bin/sqz"
elif command -v sqz >/dev/null 2>&1; then
    SQZ_BIN="sqz"
else
    SQZ_BIN=""
fi
   
if [ -n "$SQZ_BIN" ]; then
    # SQZ found: Expand to virtual context window (3x)
    VIRTUAL_CTX=$(awk -v phys="$SELECTED_CTX" 'BEGIN {print int(phys * 3.0)}')
    export AIDER_MAX_CONTEXT_WINDOW="$VIRTUAL_CTX"
    USE_SQZ="true"
else
    # SQZ not found: Fallback to physical context window
    export AIDER_MAX_CONTEXT_WINDOW="$SELECTED_CTX"
    USE_SQZ="false"
fi
   
# ==============================================================================
# 5. EXECUTE THE NATIVE FREEBSD AIDER ENGINE FLOW WITH POSIX-SAFE SQZ PIPING
# ==============================================================================
echo "🤖 Initialising Aider Client Session..."
if [ "$USE_SQZ" = "true" ]; then
    echo "📊 Context Layer   : SQZ Intelligence Active [ojuschugh1/sqz]"
    echo "📊 Hardware Buffer : ${SELECTED_CTX} tokens | Virtual Window: ${AIDER_MAX_CONTEXT_WINDOW} tokens"
    echo "────────────────────────────────────────────────────────────────────────"
    
    # Running with sqz - saving lint output in /tmp and pipe to sqz then read back exit code to aider 
    "$AIDER_BIN_EXEC" \
     --lint-cmd "c: sh -c 'clang -fsyntax-only \"\$@\" > /tmp/aider_lint.log 2>&1; RC=\$?; cat /tmp/aider_lint.log | $SQZ_BIN; exit \$RC'" \
     --test-cmd "sh -c '$TEST_RUNNER > /tmp/aider_test.log 2>&1; RC=\$?; cat /tmp/aider_test.log | $SQZ_BIN; cat /tmp/aider_test.log; exit \$RC'" \
     --git \
     --edit-format diff --no-show-model-warnings --yes-always
else
    echo "📊 Context Layer   : Native Hardware Direct (SQZ not found)"
    echo "📊 Allocated Buffer: ${AIDER_MAX_WINDOW} tokens"
    echo "────────────────────────────────────────────────────────────────────────"
    
    # No sqz found
    "$AIDER_BIN_EXEC" \
     --lint-cmd "c: clang -fsyntax-only" \
     --test-cmd "$TEST_RUNNER" \
     --git \
     --edit-format diff --no-show-model-warnings --yes-always
fi

exit 0

You can see the output in aider:
.....
cpu_stats.c
Applied edit to cpu_stats.c
Commit 781fcd1 fix: implement cpu temp from sysctl and report system load averages
[sqz] 3/3 tokens (0% reduction)
....


Another way is to run sqz stats:
$ sqz stats

📊 sqz compression stats
──────────────────────────────────────────────────

0 tokens saved
↓ 0.0% average reduction

Compressions 2
Tokens in 7
Tokens out 7
Tokens saved 0
Avg reduction 0.0%

🗄 Cache
──────────────────────────────────────────────────
Entries 1
Size 804 B



/grandpa
 
Talking to myself again.

Here's a technical aspect. It's very difficult to get claude to write to disk if you are low on NVRAM and RAM.
....

Some notes about claude. It autoupdates. That is kind of scary. I can't remember setting it to let itself be updated but I assume it can be turned off somewhere.

Anyway, autoupdate might not be so bad after all since suddenly claude started writing to disk. Of course that could also be the models that I just downloaded and are testing.

If it is model dependent then this model works fine with claude - granite-4.2-3b-Q8_0.gguf and granite-4.2-8b-Q5_K_M.gguf.

/grandpa
 
Another aspect for stubborn old fools like me trying to run AI models on hardware that has parts from 2011 ...

I still have to see a credible explanation about why that's a problem or challenge. What binary operation does it do that can't be replaced by anything else, like 2 systems of 2011 hardware working together?
GPU processing can be virtualized, distributed and simulated. I doubt the central hardware requirement of AI systems. We have flops. What's still not enough? Calculating perspectve in 3d space needs floating point operations too...
 
I still have to see a credible explanation about why that's a problem or challenge. What binary operation does it do that can't be replaced by anything else, like 2 systems of 2011 hardware working together?
GPU processing can be virtualized, distributed and simulated. I doubt the central hardware requirement of AI systems. We have flops. What's still not enough? Calculating perspectve in 3d space needs floating point operations too...

That sentence was more of a personal nature and has nothing to do with it being a problem. For a fool like me it is challenging though since I basically knew very little about how AI works not so long ago. So not a technical aspect. I like to get old things to work.

I still don't know much about how AI works even though I did compile and run a software a couple of years ago that had a model that you fed a series of pictures and the model learned to pick out the pictures that were t-shirts by setting rules.

But I think I understand your question - why does it require so much computational power? Maybe because the data has become vast, maybe because the computations are long and complicated - I don't know.

What I am seeing is that an AI model is a frozen bubble in time. It is not even aware of what time or date it is - unless of course you tell it. Maybe it's like a big bubble of matrix math? If that is the case then John Carmack could probably make it run on a 286 ...

/grandpa
 
...
Here' s a neat tool to catch-and-compress the context window before it's being sent to the model.
...

Small update on this, to make it run with claude it is required to install sqz-mcp.

$ cargo install sqz-mcp


Add .cargo/bin/ to path or edit .claude.json and add:
JSON:
    "sqz": {
      "_sqz_managed": true,
      "command": "~/.cargo/bin/sqz-mcp",
      "args": [
        "--transport",
        "stdio"
      ]
    }

/grandpa
 
That sentence was more of a personal nature and has nothing to do with it being a problem. For a fool like me it is challenging though since I basically knew very little about how AI works not so long ago. So not a technical aspect. I like to get old things to work.

I still don't know much about how AI works even though I did compile and run a software a couple of years ago that had a model that you fed a series of pictures and the model learned to pick out the pictures that were t-shirts by setting rules.

But I think I understand your question - why does it require so much computational power? Maybe because the data has become vast, maybe because the computations are long and complicated - I don't know.

What I am seeing is that an AI model is a frozen bubble in time. It is not even aware of what time or date it is - unless of course you tell it. Maybe it's like a big bubble of matrix math? If that is the case then John Carmack could probably make it run on a 286 ...

/grandpa
There should be an elementary operation that consists of a circuit of transisrtors, like all digital devices. I don't believe it has to be represented by a central module like a GPU.
What's the problem of creating that in only software-logic? Show me the component that nobody can make except AI brokers and must be made of real hardware.
Carmack probably makes a chatgpt equivalent in 486 assembly. Easy... :cool:
 
If you get tired of aider only having max 3 reflections just edit the file base_coder.py and change the value for max_reflections to whatever you want.

For my installation of aider in $HOME/venvs/aider-env the correct base_coder.py (there is more than one) was found in
$HOME/venvs/aider-env/lib/python3.12/site-packages/aider/coders/base_coder.py

/grandpa
 
Back
Top