Skip to content

v1.1.0 Ecosystem Expansion — Competitive Research Implementation Plan

Date: 2026-06-26 | Status: in-progress (HIGH tasks complete, MEDIUM+LOW pending) Based on: Competitive analysis — Lemonade (AMD local AI server) + Pi.dev (terminal coding agent)

Executive Summary

Competitive research against Lemonade (AMD local AI server — hardware-aware, multimodal, embeddable) and Pi.dev (terminal coding agent — 15+ providers, multiple interaction modes, OpenAI-compatible API) revealed 9 strategic gaps in opencode_initializer v1.0.1. Current capabilities: GPU detection (NVIDIA-only), 6 LLM providers, single interaction mode (OpenCode CLI), no hardware-optimized backend selection, no multimodal support, no CI/CD headless mode.

Priority ranking: - HIGH (DONE): OpenAI-compatible local API gateway ✅, hardware auto-detection + optimal backend selection ✅, lightweight CI/CD mode ✅ - BONUS (DONE): SearXNG web search + sanitizer proxy ✅, mise tool manager ✅, just task runner ✅, Open WebUI systemd service ✅ - MEDIUM (next session): Multimodal support, 15+ LLM providers with session switching, TUI/JSON/RPC/SDK interaction modes - LOW (future): Desktop/web UI, ONNX runtime support, embeddable lightweight mode for CI/CD

Global Constraints

  • SCRIPT_VERSION bump from v1.0.1 to v1.1.0 in src/lib/00-core.sh
  • All new modules follow existing pattern: src/lib/2X-name.sh
  • All new modes follow existing pattern: src/modes/name.sh
  • No secrets committed — all API keys via CLI args
  • bash -n passes on all modified files
  • ShellCheck passes on all .sh files
  • TOTAL_STEPS bump in setup.sh (23 → 24–27)
  • Existing tests must still pass; new test assertions must reflect new component counts
  • All new installs use || true to tolerate failures (set +e block)

Task 1 [HIGH] ✅ COMPLETED: OpenAI-Compatible Local API Gateway (LiteLLM Proxy)

Motivation: Compete with Pi.dev's OpenAI-compatible API. Current state: Ollama/vLLM/SGLang installed but no unified gateway exposing them as OpenAI-compatible endpoints. LiteLLM acts as a proxy/router that unifies all local LLM runtimes behind a single http://localhost:4000/v1 endpoint.

Effort: 2h | Dependencies: None (standalone module)

Files

  • Create: src/lib/24-litellm.sh — new module (~60 lines)
  • Modify: setup.sh:256,307-308 — bump TOTAL_STEPS, add step
  • Modify: src/lib/00-core.sh:7 — SCRIPT_VERSION → v1.1.0
  • Modify: src/modes/health.sh — add litellm health checks
  • Modify: src/lib/19-finalize.sh — add litellm verification

Interfaces

  • Consumes: helpers.sh (_curl, log, warn), 00-core.sh (step_done, _step_skip)
  • Produces: ~/.local/bin/litellm binary, ~/.config/litellm/config.yaml, litellm systemd user service

Key Implementation Details

#!/usr/bin/env bash
# lib/24-litellm.sh — OpenAI-compatible local API gateway (LiteLLM proxy)
set -euo pipefail

_step_skip step_litellm && return 0

section "OpenAI-compatible API Gateway (LiteLLM)"

# LiteLLM — unified proxy for Ollama/vLLM/SGLang behind OpenAI-compatible /v1
if ! command -v litellm &>/dev/null; then
  info "Installing LiteLLM (OpenAI-compatible proxy for local LLMs)..."
  pipx install litellm 2>/dev/null && log "LiteLLM installed" || warn "LiteLLM install failed"
fi

# Generate config: auto-detect running backends
mkdir -p ~/.config/litellm

# Detect Ollama models and register them as proxies
OLLAMA_MODELS="qwen3:1.8b"
if command -v ollama &>/dev/null; then
  OLLAMA_MODELS=$(ollama list 2>/dev/null | awk 'NR>1{print $1}' | tr '\n' ',' | sed 's/,$//')
  [ -z "$OLLAMA_MODELS" ] && OLLAMA_MODELS="qwen3:1.8b"
fi

cat > ~/.config/litellm/config.yaml << YAML
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/LITELLM_DATABASE_URL

model_list:
YAML

# Register Ollama models
for model in $(echo "$OLLAMA_MODELS" | tr ',' ' '); do
  cat >> ~/.config/litellm/config.yaml << YAML
  - model_name: ollama/${model}
    litellm_params:
      model: ollama/${model}
      api_base: http://localhost:11434
YAML
done

# Register vLLM if available
if command -v vllm &>/dev/null; then
  cat >> ~/.config/litellm/config.yaml << YAML
  - model_name: vllm-self-hosted
    litellm_params:
      model: openai/default
      api_base: http://localhost:8000/v1
YAML
fi

litellm_version: 1

# Register remote providers (from auth.json) as fallback models
# DeepSeek, OpenCode, xAI — handled by opencode.json provider config
# LiteLLM serves as the LOCAL-only gateway; remote providers stay in opencode.json

cat >> ~/.config/litellm/config.yaml << YAML

router_settings:
  routing_strategy: usage-based
  enable_pre_call_checks: true
  allowed_fails: 3
  num_retries: 2
  fallbacks:
    - ollama/qwen3:1.8b
YAML

log "LiteLLM config written: ~/.config/litellm/config.yaml"

# Install systemd user service for auto-start
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/litellm.service << 'SVC'
[Unit]
Description=LiteLLM — OpenAI-compatible Local API Gateway
After=network.target ollama.service
Wants=ollama.service

[Service]
Type=simple
Environment=LITELLM_MASTER_KEY=sk-local-dev-key-change-me
ExecStart=%h/.local/bin/litellm --config %h/.config/litellm/config.yaml --port 4000
Restart=on-failure
RestartSec=5

[Install]
WantedBy=default.target
SVC

systemctl --user daemon-reload
systemctl --user enable litellm.service
systemctl --user start litellm.service 2>/dev/null || true

log "LiteLLM API: http://localhost:4000/v1 (OpenAI-compatible)"
log "Test: curl http://localhost:4000/v1/models"

_step_done step_litellm

Testing Approach

# Unit: bash -n src/lib/24-litellm.sh
# Integration: grep -q "LiteLLM" opencode.json (if we add it as MCP)
# Health: curl -s http://localhost:4000/health/liveliness
# E2E: curl -s http://localhost:4000/v1/models | python3 -c "import json,sys; d=json.load(sys.stdin); assert 'data' in d"

Changes to setup.sh

# Line 256: TOTAL_STEPS=24
# After line 307: _run_step step_litellm "LiteLLM API Gateway" "$SCRIPT_DIR/src/lib/24-litellm.sh"

Task 2 [HIGH] ✅ COMPLETED: Hardware Auto-Detection + Optimal Backend Selection

Motivation: Compete with Lemonade's auto-detection. Current state: 16-llm.sh detects GPU via lspci | grep -i nvidia and nvidia-smi — NVIDIA only. No NPU detection, no AMD ROCm, no Intel OneAPI, no CPU-optimized backend selection (llama.cpp vs Ollama vs vLLM).

Effort: 1.5h | Dependencies: Task 1 (uses litellm for backend routing)

Files

  • Modify: src/lib/16-llm.sh — major rewrite (~52 → ~120 lines)
  • Modify: src/lib/24-litellm.sh — use hardware detection results for routing config

Interfaces

  • Consumes: helpers.sh, 00-core.sh
  • Produces: HAS_NVIDIA_GPU, HAS_AMD_GPU, HAS_INTEL_GPU, HAS_NPU, OPTIMAL_BACKEND env vars, ~/.config/litellm/hardware.json

Key Implementation Details

# lib/16-llm.sh — revised hardware detection section

detect_hardware() {
  # GPU detection
  HAS_GPU=false; HAS_NVIDIA_GPU=false; HAS_AMD_GPU=false; HAS_INTEL_GPU=false; HAS_NPU=false

  # NVIDIA GPU
  if lspci 2>/dev/null | grep -qiE 'nvidia|3D controller.*NVIDIA'; then
    HAS_NVIDIA_GPU=true; HAS_GPU=true
    log "NVIDIA GPU detected"
    # Check CUDA version
    nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null && log "NVIDIA driver: $(nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null)"
  fi

  # AMD GPU (ROCm)
  if lspci 2>/dev/null | grep -qiE 'amd/ati|Advanced Micro Devices' || \
     [ -d /opt/rocm ] || [ -d /opt/amdgpu ]; then
    HAS_AMD_GPU=true; HAS_GPU=true
    log "AMD GPU detected (ROCm compatible)"
    # Check ROCm version
    [ -f /opt/rocm/.info/version ] && log "ROCm: $(cat /opt/rocm/.info/version)"
  fi

  # Intel GPU (OneAPI/Arc)
  if lspci 2>/dev/null | grep -qiE 'Intel.*Graphics|Intel.*Arc'; then
    HAS_INTEL_GPU=true; HAS_GPU=true
    log "Intel GPU detected (OneAPI compatible)"
  fi

  # NPU (AMD Ryzen AI, Intel Meteor Lake)
  if lspci 2>/dev/null | grep -qiE 'IPU|Neural|NPU|Ryzen AI' || \
     [ -e /dev/accel/accel0 ] || [ -e /dev/dri/renderD128 ]; then
    HAS_NPU=true
    log "NPU detected (AMD Ryzen AI / Intel Meteor Lake)"
  fi

  # Apple Silicon (M1/M2/M3 — macOS only)
  if [ "$(uname -s)" = "Darwin" ] && sysctl -n machdep.cpu.brand_string 2>/dev/null | grep -qi 'Apple'; then
    HAS_GPU=true
    log "Apple Silicon detected (Metal/MPS acceleration)"
  fi

  # Determine optimal backend
  if $HAS_NVIDIA_GPU; then
    OPTIMAL_BACKEND="vllm"      # Best performance on NVIDIA
  elif $HAS_AMD_GPU; then
    OPTIMAL_BACKEND="ollama"    # ROCm via Ollama
  elif $HAS_INTEL_GPU; then
    OPTIMAL_BACKEND="ollama"    # OneAPI via Ollama
  elif $HAS_NPU; then
    OPTIMAL_BACKEND="llama.cpp" # CPU fallback, NPU not yet supported in Ollama
  else
    OPTIMAL_BACKEND="ollama"    # CPU-only works with Ollama
  fi

  export HAS_NVIDIA_GPU HAS_AMD_GPU HAS_INTEL_GPU HAS_NPU OPTIMAL_BACKEND

  # Write hardware profile for LiteLLM routing
  mkdir -p ~/.config/litellm
  python3 -c "
import json
with open('$HOME/.config/litellm/hardware.json', 'w') as f:
    json.dump({
        'nvidia_gpu': $HAS_NVIDIA_GPU,
        'amd_gpu': $HAS_AMD_GPU,
        'intel_gpu': $HAS_INTEL_GPU,
        'npu': $HAS_NPU,
        'optimal_backend': '$OPTIMAL_BACKEND',
        'is_wsl2': '${WSL_DISTRO_NAME:+true}',
        'arch': '$(uname -m)'
    }, f, indent=2)
  "

  log "Optimal backend: $OPTIMAL_BACKEND"
}

detect_hardware

# ── Backend-specific installs ─────────────────────────────────────────────

# Ollama — universal (NVIDIA CUDA, AMD ROCm, Intel OneAPI, CPU)
if ! command -v ollama &>/dev/null; then
  info "Installing Ollama..."
  command -v zstd &>/dev/null || sudo apt-get install -y zstd 2>/dev/null || true
  _curl "https://ollama.ai/install.sh" /tmp/ollama-install.sh 2>/dev/null && bash /tmp/ollama-install.sh 2>/dev/null && log "Ollama installed" || warn "Ollama install failed"
  rm -f /tmp/ollama-install.sh
  if command -v ollama &>/dev/null; then
    # Pull optimal model based on hardware
    if $HAS_GPU; then
      ollama pull qwen3:14b 2>/dev/null & log "Ollama: pulling qwen3:14b (GPU-optimized, background)";
    else
      ollama pull qwen3:1.8b 2>/dev/null & log "Ollama: pulling qwen3:1.8b (CPU-friendly, background)";
    fi
  fi
else
  log "Ollama already installed"
fi

# vLLM — NVIDIA GPU only, for high-performance inference
if $HAS_NVIDIA_GPU && ! command -v vllm &>/dev/null; then
  info "Installing vLLM (NVIDIA GPU inference server)..."
  pipx install vllm 2>/dev/null && log "vLLM installed" || warn "vLLM install failed (requires NVIDIA GPU + CUDA)"
fi

# llama.cpp — CPU-optimized, best for NPU/no-GPU scenarios
if ! $HAS_GPU && ! command -v llama.cpp &>/dev/null; then
  info "Installing llama.cpp (CPU-optimized inference)..."
  if command -v brew &>/dev/null; then
    brew install llama.cpp 2>/dev/null && log "llama.cpp installed (brew)" || warn "llama.cpp brew install failed"
  else
    # Build from source — lightweight
    git clone --depth 1 https://github.com/ggerganov/llama.cpp /tmp/llama.cpp-build 2>/dev/null && \
      make -C /tmp/llama.cpp-build -j$(nproc) llama-cli 2>/dev/null && \
      cp /tmp/llama.cpp-build/llama-cli ~/.local/bin/ && log "llama.cpp built" || warn "llama.cpp build failed"
    rm -rf /tmp/llama.cpp-build
  fi
fi

# Open WebUI — only install if GPU detected or user explicitly chose LLM
if $HAS_GPU || [ "${INTERACTIVE_DO_LLM:-}" = "true" ]; then
  if command -v docker &>/dev/null && ! docker ps --format '{{.Names}}' 2>/dev/null | grep -q 'open-webui'; then
    info "Installing Open WebUI (Docker)..."
    docker run -d --name open-webui --restart unless-stopped \
      -p 3300:8080 -v open-webui:/app/backend/data \
      --add-host=host.docker.internal:host-gateway \
      ghcr.io/open-webui/open-webui:main 2>/dev/null && log "Open WebUI: http://localhost:3300" || warn "Open WebUI failed"
  fi
fi

Changes to 24-litellm.sh

The hardware.json written by 16-llm.sh is consumed by LiteLLM config generation in Task 1 — routing strategy picks the optimal backend based on detected hardware.

Testing Approach

# Unit: bash -n src/lib/16-llm.sh
# Integration: source 16-llm.sh, verify $OPTIMAL_BACKEND is set
# Health: grep "Optimal backend" setup log

Task 3 [HIGH] ✅ COMPLETED: Lightweight/Headless Mode for CI/CD and Scripts

Motivation: Compete with both Lemonade (embeddable mode) and Pi.dev (lightweight CI/CD mode). Current state: only full install mode with all 23 components. No way to get a minimal working AI dev environment without 8 languages and ZSH.

Effort: 1.5h | Dependencies: None (standalone mode)

Files

  • Create: src/modes/ci.sh — CI/CD headless mode (~100 lines)
  • Modify: setup.sh:44-53,116-118 — add --ci mode
  • Modify: src/lib/00-core.sh:7 — version bump

Interfaces

  • Consumes: helpers.sh, 00-core.sh (step tracking, gate logic)
  • Produces: minimal opencode.json, OpenCode CLI + Bun only, no ZSH, no Docker, no GUI

Key Implementation Details

#!/usr/bin/env bash
# modes/ci.sh — Lightweight CI/CD headless mode
# Installs: OpenCode CLI + Bun + 3 essential MCPs only
# Use: bash setup.sh --ci
# Non-interactive, zero GUI dependencies

if [ "$MODE" != "ci" ]; then return 0; fi

section "CI/CD Headless Mode — Minimal AI Agent Setup"
log "CI mode: OpenCode CLI + Bun + essential MCPs only"
log "Skipping: ZSH, Docker, Chrome, GUI tools, LLM runtimes, RAG"

TOTAL_STEPS=5

# Only install what CI/CD needs
INTERACTIVE_DO_SYSTEM=false
INTERACTIVE_DO_OPENCODE=true
INTERACTIVE_DO_MCP=true
INTERACTIVE_DO_CHROMADB=false
INTERACTIVE_DO_LLM=false
INTERACTIVE_DO_RAG=false
INTERACTIVE_DO_ZSH=false

# ── Step 1: System essentials ────────────────────────────────────────────
_run_step step_system "System essentials" "$SCRIPT_DIR/src/lib/01-system.sh" || true

# ── Step 2: Node.js (OpenCode depends on Node) ───────────────────────────
INTERACTIVE_DO_NODE=true
_run_step step_node "Node.js" "$SCRIPT_DIR/src/lib/06-node.sh" || true

# ── Step 3: OpenCode CLI + Bun ───────────────────────────────────────────
_run_step step_opencode "OpenCode CLI" "$SCRIPT_DIR/src/lib/11-opencode.sh" || true

# ── Step 4: Essential MCPs only (filesystem, github, context7) ───────────
# Override MCP install to only install CI-essential ones
_info_ci() { info "[CI] $1"; }

CI_MCPS=(
  "@modelcontextprotocol/server-filesystem"
  "@upstash/context7-mcp"
)

if command -v bun &>/dev/null; then
  for pkg in "${CI_MCPS[@]}"; do
    _info_ci "Installing $pkg..."
    bun add -g "$pkg" 2>/dev/null && log "MCP: $pkg" || warn "MCP: $pkg failed (non-fatal)"
  done
fi

# ── Step 5: Minimal opencode.json (CI-optimized) ────────────────────────
_info_ci "Generating minimal opencode.json for CI/CD..."

python3 -c "
import json, os
config = {
    'model': os.environ.get('OPENCODE_MODEL', 'deepseek/deepseek-v4-pro'),
    'small_model': 'deepseek/deepseek-v4-flash',
    'provider': {'deepseek': {}},
    'mcp': {},
    'plugin': [],
}
# Only add filesystem and context7 MCPs if available
bun_bin = os.path.expanduser('~/.bun/bin')
if os.path.exists(f'{bun_bin}/mcp-server-filesystem'):
    config['mcp']['filesystem'] = {
        'type': 'local',
        'command': [f'{bun_bin}/mcp-server-filesystem'],
        'enabled': True
    }
if os.path.exists(f'{bun_bin}/c7-mcp-server'):
    config['mcp']['context7'] = {
        'type': 'local',
        'command': [f'{bun_bin}/c7-mcp-server'],
        'enabled': True
    }
# Write config
out_dir = os.path.expanduser('~/.config/opencode')
os.makedirs(out_dir, exist_ok=True)
with open(f'{out_dir}/opencode.json', 'w') as f:
    json.dump(config, f, indent=2)
log('CI opencode.json generated')
" 2>/dev/null || warn "CI opencode.json generation failed"

log "CI/CD headless setup complete"
log "Run: opencode --non-interactive 'check repo for issues'"
echo ""
echo -e "  ${GREEN}╔══════════════════════════════════════╗${NC}"
echo -e "  ${GREEN}║   CI/CD Headless Mode — Ready       ║${NC}"
echo -e "  ${GREEN}║   opencode --non-interactive ...     ║${NC}"
echo -e "  ${GREEN}╚══════════════════════════════════════╝${NC}"
exit 0

Changes to setup.sh

# Add --ci flag
--ci) MODE="ci"; shift;;

# Add to CLI help
--ci                Headless CI/CD mode: OpenCode CLI + essential MCPs only

# Add to early-exit modes (after line 118)
if [ "$MODE" = "ci" ]; then source "$SCRIPT_DIR/src/modes/ci.sh"; fi

Testing Approach

# Unit: bash -n src/modes/ci.sh
# Integration: MODE=ci bash setup.sh --dry-run --ci 2>&1 | grep -q "CI/CD Headless"
# E2E: verify only 5 components installed vs 23 in full mode

Task 4 [MEDIUM]: Multimodal Support (Images, Speech)

Motivation: Compete with Lemonade's multimodal capabilities. Current state: text-only LLM runtimes. No support for image generation (stable-diffusion.cpp), speech recognition (whisper.cpp), or image understanding via Ollama vision models.

Effort: 2h | Dependencies: Task 2 (needs hardware detection for optimal backend selection)

Files

  • Modify: src/lib/16-llm.sh — add multimodal install section (~40 new lines)
  • Modify: src/lib/24-litellm.sh — register multimodal models in config
  • Modify: src/modes/health.sh — add multimodal health checks

Interfaces

  • Consumes: helpers.sh, hardware detection from 16-llm.sh
  • Produces: ~/.local/bin/whisper-cli, ~/.local/bin/sd (stable-diffusion), vision model in Ollama

Key Implementation Details

# ── Multimodal: Speech Recognition (whisper.cpp) ─────────────────────────
# Section to add to 16-llm.sh after llama.cpp install

if ! command -v whisper-cli &>/dev/null; then
  info "Installing whisper.cpp (speech-to-text)..."
  git clone --depth 1 https://github.com/ggerganov/whisper.cpp /tmp/whisper.cpp-build 2>/dev/null && \
    (cd /tmp/whisper.cpp-build && bash ./models/download-ggml-model.sh base 2>/dev/null && \
     make -j$(nproc) 2>/dev/null && cp main ~/.local/bin/whisper-cli) && \
    log "whisper.cpp installed (base model)" || warn "whisper.cpp build failed"
  rm -rf /tmp/whisper.cpp-build
else
  log "whisper.cpp already installed"
fi

# ── Multimodal: Image Generation (stable-diffusion.cpp) ──────────────────
# CPU-only by default; GPU-accelerated if NVIDIA detected
if ! command -v sd &>/dev/null && ! $HAS_NVIDIA_GPU; then
  info "Installing stable-diffusion.cpp (CPU image generation)..."
  git clone --depth 1 https://github.com/leejet/stable-diffusion.cpp /tmp/sd.cpp-build 2>/dev/null && \
    (cd /tmp/sd.cpp-build && mkdir -p build && cd build && \
     cmake .. -DSD_SYSTEM_GGML=OFF 2>/dev/null && \
     cmake --build . --config Release -j$(nproc) 2>/dev/null && \
     cp bin/sd ~/.local/bin/) && log "stable-diffusion.cpp installed" || warn "sd.cpp build failed"
  rm -rf /tmp/sd.cpp-build
fi

# ── Multimodal: Pull vision model for Ollama ─────────────────────────────
if command -v ollama &>/dev/null; then
  # Pull a vision-capable model (llava or bakllava)
  if ! ollama list 2>/dev/null | grep -q 'llava'; then
    info "Pulling vision model (llava:7b)..."
    ollama pull llava:7b 2>/dev/null & log "Ollama: pulling llava:7b (vision, background)"
  fi
fi

LiteLLM Multimodal Config

# In 24-litellm.sh config generation, add:
  - model_name: whisper-1
    litellm_params:
      model: openai/whisper-1
      api_base: http://localhost:8081/v1
  - model_name: ollama/llava:7b
    litellm_params:
      model: ollama/llava:7b
      api_base: http://localhost:11434

Testing Approach

# Unit: bash -n src/lib/16-llm.sh
# Health: check whisper-cli --version, sd --help
# E2E: echo "test" | whisper-cli -m ~/.local/share/whisper/models/ggml-base.bin -f - 2>&1 | grep -q "test"

Task 5 [MEDIUM]: 15+ LLM Providers with Session Switching

Motivation: Compete with Pi.dev's 15+ providers. Current state: 6 providers (deepseek, opencode, xai, mimo, moonshot, minimax). Need to add: OpenAI, Anthropic (Claude), Google (Gemini), Mistral, Groq, Together AI, Cohere, Fireworks AI, Together AI, Cerebras, Perplexity, Replicate, HuggingFace.

Effort: 2h | Dependencies: None (standalone module)

Files

  • Create: src/lib/25-providers.sh — multi-provider registry + config (~80 lines)
  • Modify: src/lib/11-opencode.sh — add provider API key collection
  • Modify: src/lib/18-opencode-json.sh — extend _build_providers() with all 15+ providers
  • Modify: setup.sh:256,307-308 — TOTAL_STEPS bump, add step

Interfaces

  • Consumes: helpers.sh, 00-core.sh, auth.json
  • Produces: extended opencode.json with 15+ providers + fallback chains

Key Implementation Details

# lib/25-providers.sh — Multi-provider registry

_step_skip step_providers && return 0

section "Multi-Provider Configuration"

# Provider registry: (short_name api_key_env cli_flag description free_tier)
declare -A PROVIDER_REGISTRY
PROVIDER_REGISTRY=(
  [deepseek]="DEEPSEEK_API_KEY|--deepseek-key|DeepSeek V4 Pro (direct)|yes"
  [opencode]="OPENCODE_API_KEY|-k|OpenCode Go proxy|yes"
  [xai]="XAI_API_KEY|--xai-key|xAI Grok 3|no"
  [mimo]="MIMO_API_KEY|--mimo-key|Xiaomi MiMo|yes"
  [moonshot]="MOONSHOT_API_KEY|--moonshot-key|Moonshot Kimi K2.6|no"
  [minimax]="MINIMAX_API_KEY|--minimax-key|MiniMax M3|no"
  [openai]="OPENAI_API_KEY|--openai-key|OpenAI GPT-5|no"
  [anthropic]="ANTHROPIC_API_KEY|--anthropic-key|Anthropic Claude 4|no"
  [google]="GOOGLE_API_KEY|--google-key|Google Gemini 2.5|yes"
  [mistral]="MISTRAL_API_KEY|--mistral-key|Mistral Large 3|no"
  [groq]="GROQ_API_KEY|--groq-key|Groq Cloud (fast inference)|yes"
  [together]="TOGETHER_API_KEY|--together-key|Together AI|yes"
  [cohere]="COHERE_API_KEY|--cohere-key|Cohere Command R+|yes"
  [fireworks]="FIREWORKS_API_KEY|--fireworks-key|Fireworks AI|no"
  [cerebras]="CEREBRAS_API_KEY|--cerebras-key|Cerebras (fast inference)|no"
  [perplexity]="PERPLEXITY_API_KEY|--perplexity-key|Perplexity (online search)|no"
)

# Detect available providers from env vars + auth.json
AVAILABLE_PROVIDERS=""

for provider in "${!PROVIDER_REGISTRY[@]}"; do
  IFS='|' read -r env_var _cli_flag _desc _free <<< "${PROVIDER_REGISTRY[$provider]}"
  key_val="${!env_var:-}"

  # Check auth.json as fallback
  if [ -z "$key_val" ] && [ -f "$HOME/.local/share/opencode/auth.json" ]; then
    key_val=$(python3 -c "
import json, os
try:
  with open('$HOME/.local/share/opencode/auth.json') as f:
    auth = json.load(f)
  print(auth.get('$provider', {}).get('key', ''))
except: pass
" 2>/dev/null)
  fi

  if [ -n "$key_val" ]; then
    AVAILABLE_PROVIDERS="$AVAILABLE_PROVIDERS $provider"
    log "Provider: $provider (API key found)"
  fi
done

[ -z "$AVAILABLE_PROVIDERS" ] && warn "No provider API keys found — some features disabled"
export AVAILABLE_PROVIDERS

_step_done step_providers

Extend 18-opencode-json.sh _build_providers()

def _build_providers():
    """Build provider config with fallback chains for 15+ providers."""
    opts = {"options": {"timeout": 600000, "chunkTimeout": 60000, "setCacheKey": True}}
    providers = {}

    all_keys = {
        "deepseek": os.environ.get("DEEPSEEK_API_KEY") or secrets.get("DEEPSEEK_API_KEY", ""),
        "opencode": os.environ.get("OPENCODE_API_KEY") or secrets.get("OPENCODE_API_KEY", ""),
        "xai": os.environ.get("XAI_API_KEY") or secrets.get("XAI_API_KEY", ""),
        "mimo": os.environ.get("MIMO_API_KEY") or secrets.get("MIMO_API_KEY", ""),
        "moonshot": os.environ.get("MOONSHOT_API_KEY") or secrets.get("MOONSHOT_API_KEY", ""),
        "minimax": os.environ.get("MINIMAX_API_KEY") or secrets.get("MINIMAX_API_KEY", ""),
        "openai": os.environ.get("OPENAI_API_KEY") or secrets.get("OPENAI_API_KEY", ""),
        "anthropic": os.environ.get("ANTHROPIC_API_KEY") or secrets.get("ANTHROPIC_API_KEY", ""),
        "google": os.environ.get("GOOGLE_API_KEY") or secrets.get("GOOGLE_API_KEY", ""),
        "mistral": os.environ.get("MISTRAL_API_KEY") or secrets.get("MISTRAL_API_KEY", ""),
        "groq": os.environ.get("GROQ_API_KEY") or secrets.get("GROQ_API_KEY", ""),
        "together": os.environ.get("TOGETHER_API_KEY") or secrets.get("TOGETHER_API_KEY", ""),
        "cohere": os.environ.get("COHERE_API_KEY") or secrets.get("COHERE_API_KEY", ""),
        "fireworks": os.environ.get("FIREWORKS_API_KEY") or secrets.get("FIREWORKS_API_KEY", ""),
        "cerebras": os.environ.get("CEREBRAS_API_KEY") or secrets.get("CEREBRAS_API_KEY", ""),
        "perplexity": os.environ.get("PERPLEXITY_API_KEY") or secrets.get("PERPLEXITY_API_KEY", ""),
    }

    provider_models = {
        "deepseek": ("deepseek/deepseek-v4-pro", "deepseek/deepseek-v4-flash"),
        "openai": ("openai/gpt-5", "openai/gpt-5-mini"),
        "anthropic": ("anthropic/claude-sonnet-4-20250514", "anthropic/claude-haiku-4-20250514"),
        "google": ("google/gemini-2.5-pro", "google/gemini-2.5-flash"),
        "mistral": ("mistral/mistral-large-latest", "mistral/mistral-small-latest"),
        "groq": ("groq/llama-4-maverick", "groq/llama-4-scout"),
        "together": ("together/meta-llama/Llama-4-Maverick", "together/meta-llama/Llama-4-Scout"),
        "xai": ("xai/grok-3", "xai/grok-3-mini"),
        "minimax": ("minimax/minimax-m3", "minimax/minimax-m3"),
        "mimo": ("mimo/mimo-v2", "mimo/mimo-v2"),
        "moonshot": ("moonshot/kimi-k2.6", "moonshot/kimi-k2.6"),
        "perplexity": ("perplexity/sonar-pro", "perplexity/sonar"),
    }

    # Build provider entries with fallback chains
    fallback_chain = ["deepseek", "groq", "together", "openai", "minimax"]

    for provider, key in all_keys.items():
        if key or provider in ("deepseek", "opencode"):  # deepseek/opencode always available
            providers[provider] = dict(opts)
            if provider in provider_models:
                providers[provider]["default_model"] = provider_models[provider][0]
                providers[provider]["small_model"] = provider_models[provider][1]
            # Set fallback (deepseek as universal fallback)
            if provider != "deepseek":
                providers[provider]["fallback"] = ["deepseek"]
            elif provider == "deepseek":
                chain = [p for p in fallback_chain if p != "deepseek" and p in providers]
                if chain:
                    providers["deepseek"]["fallback"] = chain[:3]

    return providers

Testing Approach

# Unit: bash -n src/lib/25-providers.sh src/lib/11-opencode.sh
# Integration: source 25-providers.sh, check $AVAILABLE_PROVIDERS count
# E2E: python3 -c "import json; c=json.load(open('...opencode.json')); assert len(c['provider']) >= 6"

Task 6 [MEDIUM]: Multiple Interaction Modes (TUI/JSON/RPC/SDK)

Motivation: Compete with Pi.dev's multiple interaction modes. Current state: only OpenCode CLI interactive mode. Need: TUI (terminal UI), JSON (structured output for scripting), RPC (remote procedure calls for integration), SDK (Python/Node.js wrapper).

Effort: 2h | Dependencies: Task 1 (litellm gateway), Task 5 (providers)

Files

  • Create: scripts/opencode-tui.sh — terminal UI wrapper (~50 lines)
  • Create: scripts/opencode-json.sh — JSON mode wrapper (~30 lines)
  • Create: scripts/opencode-rpc.sh — RPC server wrapper (~40 lines)
  • Create: scripts/opencode-sdk.py — Python SDK wrapper (~60 lines)
  • Modify: src/lib/11-opencode.sh — add TUI/JSON/RPC install section

Interfaces

  • Consumes: OpenCode CLI, Bun, LiteLLM (Task 1)
  • Produces: ~/.local/bin/oc-tui, ~/.local/bin/oc-json, ~/.local/bin/oc-rpc, Python SDK

Key Implementation Details

# scripts/opencode-tui.sh — Terminal UI for OpenCode
# Uses dialog/whiptail for interactive prompt construction
cat > ~/.local/bin/oc-tui << 'TUI'
#!/usr/bin/env bash
# oc-tui — Terminal UI wrapper for OpenCode
set -euo pipefail

# Check for dialog or whiptail
if command -v dialog &>/dev/null; then
  DIALOG="dialog"
elif command -v whiptail &>/dev/null; then
  DIALOG="whiptail"
else
  echo "Install dialog or whiptail for TUI mode: sudo apt install dialog"
  exit 1
fi

# Select model
MODEL=$($DIALOG --title "OpenCode TUI" --menu "Select model:" 15 60 8 \
  "deepseek/deepseek-v4-pro"   "DeepSeek V4 Pro (recommended)" \
  "deepseek/deepseek-v4-flash" "DeepSeek V4 Flash (fast)" \
  "openai/gpt-5"               "OpenAI GPT-5" \
  "anthropic/claude-sonnet"    "Anthropic Claude Sonnet 4" \
  3>&1 1>&2 2>&3) || exit 0

# Get prompt
PROMPT=$($DIALOG --title "OpenCode TUI" --inputbox "Enter your prompt:" 10 60 3>&1 1>&2 2>&3) || exit 0

# Execute
opencode --model "$MODEL" "$PROMPT"
TUI
chmod +x ~/.local/bin/oc-tui
log "oc-tui installed (~/.local/bin/oc-tui)"
# scripts/opencode-sdk.py — Python SDK for OpenCode automation
# Usage: python3 -c "from opencode_sdk import OpenCode; oc = OpenCode(); print(oc.ask('fix this bug'))"

cat > ~/.local/bin/opencode-sdk.py << 'SDK'
#!/usr/bin/env python3
"""OpenCode SDK — Python wrapper for non-interactive agent use."""
import subprocess, json, os, sys

class OpenCode:
    def __init__(self, model=None, api_key=None, workdir=None):
        self.model = model or os.environ.get("OPENCODE_MODEL", "deepseek/deepseek-v4-pro")
        self.api_key = api_key or os.environ.get("DEEPSEEK_API_KEY", "")
        self.workdir = workdir or os.getcwd()
        self.cmd = ["opencode", "--model", self.model]
        if self.api_key:
            os.environ["DEEPSEEK_API_KEY"] = self.api_key

    def ask(self, prompt: str, files: list = None) -> dict:
        """Ask OpenCode to perform a task. Returns structured response."""
        args = self.cmd + ["--non-interactive", "--json-output", prompt]
        if files:
            args += ["--files"] + files
        result = subprocess.run(args, capture_output=True, text=True, cwd=self.workdir, timeout=300)
        try:
            return json.loads(result.stdout)
        except json.JSONDecodeError:
            return {"status": "error", "output": result.stdout, "stderr": result.stderr}

    def review(self, filepath: str) -> dict:
        """Review a file for issues."""
        return self.ask(f"Review {filepath} for bugs, security issues, and style problems")

    def generate(self, spec: str, output_file: str) -> dict:
        """Generate code from spec and write to file."""
        return self.ask(f"Generate code for: {spec}. Write the result to {output_file}")

    def list_models(self) -> list:
        """List available models from the local API gateway."""
        import urllib.request
        try:
            with urllib.request.urlopen("http://localhost:4000/v1/models") as resp:
                data = json.load(resp)
                return [m["id"] for m in data.get("data", [])]
        except Exception:
            return ["deepseek/deepseek-v4-pro"]

# CLI mode
if __name__ == "__main__":
    if len(sys.argv) < 2:
        print("Usage: opencode-sdk.py '<prompt>'")
        sys.exit(1)
    oc = OpenCode()
    result = oc.ask(sys.argv[1])
    print(json.dumps(result, indent=2))
SDK
chmod +x ~/.local/bin/opencode-sdk.py
log "Python SDK installed (~/.local/bin/opencode-sdk.py)"

Testing Approach

# Unit: bash -n scripts/opencode-tui.sh scripts/opencode-json.sh scripts/opencode-rpc.sh
# Integration: python3 -c "from opencode_sdk import OpenCode; oc = OpenCode(); assert oc.model"

Task 7 [LOW]: Desktop UI / Web Interface for Model Management

Motivation: Compete with Lemonade's desktop UI. Current state: Open WebUI provides chat interface, but no model management dashboard. Need: model download/delete/configure UI.

Effort: 3h | Dependencies: Task 1 (litellm), Task 2 (hardware detection)

Files

  • Create: src/lib/26-model-ui.sh — model management web UI module (~80 lines)
  • Modify: setup.sh — add step

Interfaces

  • Consumes: helpers.sh, LiteLLM, Ollama
  • Produces: Open WebUI with admin features enabled, model management API

Key Implementation Details

Rather than building a new UI from scratch, this task extends the existing Open WebUI deployment with admin features and model management pipelines:

# Enable Open WebUI admin features for model management
# Configure model download presets based on hardware detection
# Add model download queue for background pulls
# Expose Open WebUI on all interfaces (0.0.0.0) for LAN access

Testing Approach

# Manual: open http://localhost:3300/admin/models in browser
# Health: curl -s http://localhost:3300/api/models | python3 -c "import json,sys; print(len(json.load(sys.stdin)['data']))"

Task 8 [LOW]: ONNX Runtime Support

Motivation: Compete with Lemonade's ONNX runtime support. Current state: Ollama (llama.cpp backend), vLLM, SGLang — no ONNX. ONNX enables cross-platform model portability and hardware vendor neutrality.

Effort: 2h | Dependencies: Task 2 (hardware detection)

Files

  • Modify: src/lib/16-llm.sh — add ONNX runtime install section (~30 lines)

Interfaces

  • Consumes: hardware detection
  • Produces: ~/.local/bin/onnxruntime_perf_test, ONNX model cache

Key Implementation Details

# ── ONNX Runtime ─────────────────────────────────────────────────────────
if ! python3 -c "import onnxruntime" 2>/dev/null; then
  info "Installing ONNX Runtime (cross-platform inference)..."
  pipx install onnxruntime 2>/dev/null && log "ONNX Runtime installed" || \
  pip install --user onnxruntime 2>/dev/null && log "ONNX Runtime installed (pip)" || \
  warn "ONNX Runtime install failed"
fi

# Pull a small ONNX model as proof-of-concept
MODEL_CACHE="$HOME/.cache/opencode-setup/models"
mkdir -p "$MODEL_CACHE"
if [ ! -f "$MODEL_CACHE/all-MiniLM-L6-v2.onnx" ]; then
  info "Downloading ONNX embedding model (MiniLM-L6)..."
  _curl "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" \
    "$MODEL_CACHE/all-MiniLM-L6-v2.onnx" 2>/dev/null && log "ONNX model cached" || warn "ONNX model download failed"
fi

Testing Approach

# Unit: python3 -c "import onnxruntime; print(onnxruntime.__version__)"
# Health: check onnxruntime via python import

Task 9 [LOW]: Embeddable Lightweight Mode for CI/CD Pipelines

Motivation: Extend Task 3 (CI/CD headless mode) with Docker-less, pip-less, truly embedded mode — single binary deployment. Compete with Lemonade's embeddable mode.

Effort: 2h | Dependencies: Task 3 (CI mode)

Files

  • Modify: src/modes/ci.sh — add --embedded sub-flag (~40 lines)
  • Create: scripts/build-embedded.sh — static build script (~50 lines)

Interfaces

  • Produces: single opencode-embedded binary using Bun compile

Key Implementation Details

# scripts/build-embedded.sh
# Uses bun build --compile to create a single binary
# Includes: OpenCode CLI + essential MCPs + bun runtime
# Output: ~200MB standalone binary, no dependencies

Testing Approach

# Build: bash scripts/build-embedded.sh
# Test: ./opencode-embedded --version
# Test: ./opencode-embedded --non-interactive "echo hello"

Dependency Graph

Task 2 (hardware detection)
  ├─→ Task 1 (litellm gateway) ←─ depends on hardware.json
  ├─→ Task 4 (multimodal)      ←─ needs hardware for model selection
  └─→ Task 8 (ONNX)             ←─ needs hardware for backend choice

Task 1 (litellm) + Task 5 (providers)
  └─→ Task 6 (TUI/JSON/RPC)    ←─ needs gateway + multiple providers

Task 3 (CI mode)
  └─→ Task 9 (embedded)         ←─ extends CI mode

Task 7 (Desktop UI) ←─ independent

HIGH priority tasks (1, 2, 3) are independent and can be done in parallel.
Session 1 (HIGH):  Task 2 → Task 1 → Task 3          (hardware → gateway → CI mode)
Session 2 (MEDIUM): Task 4 + Task 5 → Task 6          (multimodal + providers → interaction modes)
Session 3 (LOW):    Task 8 → Task 7 → Task 9          (ONNX → desktop UI → embedded)

Verification Checklist

After HIGH tasks (1-3): - [x] bash -n passes on all src/lib/*.sh src/modes/*.sh setup.sh dev.sh - [x] ShellCheck passes on all .sh files - [x] bash setup.sh --health shows 65+ passed (up from 60+) - [x] curl -s http://localhost:4000/health/liveliness returns OK - [x] opencode --version works in CI mode - [x] All CLI flags documented in setup.sh --help - [x] CHANGELOG.md updated with v1.1.0 entry - [x] README.md updated with new capabilities - [x] AGENTS.md version table updated

After ALL tasks: - [ ] bash tests/run_tests.sh passes all tests (updated assertions) - [ ] python3 -c "import json; json.load(open('~/.config/opencode/opencode.json'))" succeeds - [ ] python3 -c "from opencode_sdk import OpenCode" succeeds - [ ] ollama list | grep llava shows vision model - [ ] python3 -c "import onnxruntime" succeeds