caharkness.com

Home · Contact · Git


ROCm 10.0 Software Script

Here's a script for your Debian box that has ROCm-compatible hardware. You'll need to edit the script so that it suits you, but it gets llama.cpp and ComfyUI up and running pretty quickly.

#!/bin/bash

export STARTING_DIR="$(dirname "$0")"

cd "$STARTING_DIR"

# gfx1201 = Radeon 9070 XT and R9700
# gfx1151 = Strix Halo APUs
export AMD_GPU="gfx1201"

function install_rocm()
{
    apt update

    # Add the ROCm 10.0 official repo:
    cat << EOF > /etc/apt/sources.list.d/amdrocm-stable.sources
X-Repo-Id: amdrocm-stable
Types: deb
URIs: https://stable.repo.amd.com/rocm/core/packages/debian13/
Suites: stable
Components: main
Architectures: amd64
Enabled: yes
Signed-By:
$(wget https://stable.repo.amd.com/rocm/gpg/packages.gpg -O - | sed 's/^/ /')
EOF

    # Install the required dependencies for a ROCm/HIP build:
    apt install -y amdrocm-core-sdk10.0-${AMD_GPU}
    apt install -y amdrocm-core-dev10.0-${AMD_GPU}

}

function check_rocm()
{
    if [[ ! -d /opt/rocm/core-10.0 ]]
    then
        echo "/opt/rocm does not exist"
        echo "Was ROCm 10 installed correctly?"
        exit 1
    fi
}

export LD_LIBRARY_PATH=/opt/rocm/core-10.0/lib:$LD_LIBRARY_PATH
export LD_LIBRARY_PATH=/opt/rocm/core-10.0/lib/llvm/lib:$LD_LIBRARY_PATH

if [[ "$@" =~ "--llamacpp" ]]
then

    if [[ "$@" =~ "--build" ]]
    then
        install_rocm

        # Install the common build dependencies:
        apt install build-essential ca-certificates cmake curl git libcurl4-openssl-dev libdrm-dev libdw-dev libelf-dev libgl-dev libnuma-dev libpciaccess-dev libssl-dev libudev-dev libzstd-dev ninja-build pciutils pkg-config python3 python3-pip python3-venv xxd zlib1g-dev

        # Clone llama.cpp
        git clone https://github.com/ggml-org/llama.cpp.git

        cd llama.cpp

        # Make sure it's up to date:
        git pull

        # Create a "build" configuration for ROCm:
        cmake -B build \
            -DGGML_HIP=ON \
            -DAMD_GPU_TARGETS=${AMD_GPU} \
            -DCMAKE_BUILD_TYPE=Release

        # Actually build the "build" ROCm configuration (using 8 threads):
        cmake --build build --config Release --parallel 8

        cd ..
        exit 0
    fi

    check_rocm

    export GGML_CUDA_DISABLE_GRAPHS=1
    export ROCBLAS_INTERNAL_GEMM_CACHE_SIZE=0

    ./llama.cpp/build/bin/llama-server \
        --n-gpu-layers -1 \
        --host 0.0.0.0 \
        --port 9931 \
        --ctx-size 65536 \
        --alias "qwen3.8-27b" \
        --model "/shared/Apps/llama.cpp/models/Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-RCO-IQ3_XXS.gguf" \
        --mmproj "/shared/Apps/llama.cpp/models/mmproj-Swift-1.5-Qwen3.8-27B-Uncensored-Dynamic-MTP-RCO-BF16.gguf" \
        --batch-size 2048 \
        --ubatch-size 1024 \
        --cache-type-k q4_0 \
        --cache-type-v q4_0 \
        --jinja \
        --webui-mcp-proxy \
        --reasoning on \
        --webui-config-file "config.json" \
        --flash-attn on \
        --fit off \
        --load-mode none \
        --spec-type draft-mtp \
        --spec-draft-n-max 2 \
        --spec-draft-p-min 0.1 \
        --spec-draft-device ROCm1 \
        --split-mode layer \
        --tensor-split 1,1

    exit 0
fi

if [[ "$@" =~ "--comfyui" ]]
then

    if [[ "$@" =~ "--install" ]]
    then
        install_rocm

        git clone https://github.com/Comfy-Org/ComfyUI.git

        cd ComfyUI

        git pull

        rm -rf venv
        python3 -m venv venv

        source ./venv/bin/activate

        pip install --index-url https://stable.repo.amd.com/rocm/whl-next/ \
                "torch[device-${AMD_GPU}]==2.13.0+rocm10.0.0" \
                "torchvision[device-${AMD_GPU}]==0.28.0+rocm10.0.0" \
                "torchaudio==2.11.0.2+rocm10.0.0"

        pip install -r requirements.txt

        cd ..        
        exit 0
    fi

    check_rocm

    if [[ ! -d ./ComfyUI/venv ]]
    then
        echo "./ComfyUI/venv was not found"
        echo "Please run this script with --comfyui --install first"
        exit 1
    fi

    cd ComfyUI

    source ./venv/bin/activate

    python -u main.py \
        --listen \
        --port 8188 \
        --gpu-only \
        --disable-smart-memory \
        --disable-mmap \
        --enable-manager \
        --user-directory user

    exit 0
fi

echo "You must run this script with --llamacpp or --comfyui"
exit 1

Samba smb.conf Example

For home NAS setups running some flavor of Linux and a few Windows clients on the network, I find myself reaching for the same few lines every time, so I am providing an example below:

[nobody]
    browseable = no

[storage]
    path = /storage
    browseable = yes
    writeable = yes
    valid users = root
    force user = root
    force group = root
    create mask = 0777
    directory mask = 0777
    ;guest ok = yes

[proxmox]
    path = /storage/Proxmox
    browseable = yes
    writeable = yes
    valid users = proxmox
    force user = root
    force group = root
    create mask = 0777
    directory mask = 0777
    ;guest ok = yes

[shared]
   path = /storage/Shared
   browseable = yes
   writeable = yes
   force user = root
   force group = root
   create mask = 0777
   directory mask = 0777
   guest ok = yes

Don't forget to create Samba users with commands like smbpasswd -a root or smbpasswd -a myuser.


Windows in Docker

I lied. This is a script I use to get Windows 11 running in a Podman container. Daemonize it however you'd like!

podman run \
    --interactive \
    --tty \
    --rm \
    --name windows \
    --env "VERSION=11" \
    --publish 8006:8006 \
    --publish 3390:3389/tcp \
    --publish 3390:3389/udp \
    --device=/dev/kvm \
    --device=/dev/net/tun \
    --cap-add NET_ADMIN \
    --volume "./windows:/storage" \
    --stop-timeout 120 \
    docker.io/dockurr/windows

Activation via PowerShell:

irm https://get.activated.win | iex

Source


Debloat round 1:

& ([scriptblock]::Create((irm "https://debloat.raphi.re/")))

Source


Debloat round 2:

irm https://christitus.com/win | iex

Source


Python Setup

I work between both Windows and Linux and find myself writing the same shell script for every project I spin up. Here's now it's done:

#!/bin/bash

cd "$(dirname "$0")"

CREATE_VENV=0

if [[ ! -d venv ]]
then
    python3 -m venv venv || python -m venv venv || (
        echo "Could not create a Python virtual environment."
        exit 1
    )

    CREATE_VENV=1
fi

source ./venv/bin/activate
source ./venv/Scripts/activate

if [[ $CREATE_VENV -eq 1 ]]
then
    pip install dependency1 dependency2 etc
fi

python -u app.py $@

Keep in mind this is a shell script for Bash. To run on Windows, you must also have Bash on path. Once you do, this script works great and is portable between the big three: Windows, macOS, and Linux.


PM2 Setup

Here's my preferred way to run a script like a service. First, to install PM2 on a Debian-based system, here's what I would run as root:

apt install nodejs
apt install npm
npm install -g n
n latest
npm install -g pm2
pm2 startup

Then, in what ever directory contains the software project that needs a "run this forever" treatment, I simply run:

pm2 start script.sh --name script -- argument1 argument2
pm2 save

That will start script.sh with the name "script" and start it with all the command line arguments that come after the last two hyphens.


LLMs

Screw the frontier models.

Get yourself a llama.cpp release that's appropriate for your computer. If you have an NVIDIA GPU, be sure to download the .dll files as well and extract them to the same directory. Vulkan builds should run everywhere.

Get yourself a Qwen3.6 35B A3B quant that fits well on your GPU or in your system's RAM. This model's Q4_K_M quant replaced Qwen3-Coder-Next's Q8 quant for me. Don't forget the mmproj file to go along with it if you want OCR capabilities!

Make yourself a launch script like so:

#!/bin/bash

./llama-server \
    --n-gpu-layers -1 \
    --host 0.0.0.0 \
    --port 8080 \
    --ctx-size 262144 \
    --n-gpu-layers -1 \
    --model "Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf" \
    --mmproj "mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" \
    --ubatch-size 1024 \
    --batch-size 1024 \
    --jinja \
    --webui-mcp-proxy \
    --chat-template-kwargs "{ \"enable_thinking\": false }"

Modify it so that it points to your model's two .gguf files and run. I run this exact script on my Framework Desktop with a Ryzen AI Max+ 395 with 128 GB of unified memory and I am getting close to 70 t/s in generation with small context.


"Impose" by Bad Omens

Give it a listen and read the lyrics along with it.


Fine-tune xrdp

Set the kernel option net.core.wmem_max to 8388608 via issuing sysctl -w net.core.wmem_max=8388608. You can have this change persist across reboots by writing:

net.core.wmem_max = 8388608

...to /etc/sysctl.d/xrdp.conf.

Configure xrdp.ini

Ensure the following:

tcp_send_buffer_bytes=4194304
max_bpp=16

This was all that was needed to get my unbearably slow experience to something usable for connections over a private VPN. No tweaks to the compositor needed.

Source


Multi-session User in xrdp

Insert the following line

export $(dbus-launch)

...before the test -x and exec lines in /etc/xrdp/startwm.sh to allow signing into your user account more than once. This is especially useful for users coming from Windows environments where you can be signed in locally (at home) and connect to your same user account from afar via RDP.

Source