v1.1.0 Ecosystem Expansion — Competitive Research Implementation Plan¶
Date: 2026-06-26 | Status: in-progress (HIGH tasks complete, MEDIUM+LOW pending) Based on: Competitive analysis — Lemonade (AMD local AI server) + Pi.dev (terminal coding agent)
Executive Summary¶
Competitive research against Lemonade (AMD local AI server — hardware-aware, multimodal, embeddable) and Pi.dev (terminal coding agent — 15+ providers, multiple interaction modes, OpenAI-compatible API) revealed 9 strategic gaps in opencode_initializer v1.0.1. Current capabilities: GPU detection (NVIDIA-only), 6 LLM providers, single interaction mode (OpenCode CLI), no hardware-optimized backend selection, no multimodal support, no CI/CD headless mode.
Priority ranking: - HIGH (DONE): OpenAI-compatible local API gateway ✅, hardware auto-detection + optimal backend selection ✅, lightweight CI/CD mode ✅ - BONUS (DONE): SearXNG web search + sanitizer proxy ✅, mise tool manager ✅, just task runner ✅, Open WebUI systemd service ✅ - MEDIUM (next session): Multimodal support, 15+ LLM providers with session switching, TUI/JSON/RPC/SDK interaction modes - LOW (future): Desktop/web UI, ONNX runtime support, embeddable lightweight mode for CI/CD
Global Constraints¶
SCRIPT_VERSIONbump fromv1.0.1tov1.1.0insrc/lib/00-core.sh- All new modules follow existing pattern:
src/lib/2X-name.sh - All new modes follow existing pattern:
src/modes/name.sh - No secrets committed — all API keys via CLI args
bash -npasses on all modified files- ShellCheck passes on all .sh files
TOTAL_STEPSbump insetup.sh(23 → 24–27)- Existing tests must still pass; new test assertions must reflect new component counts
- All new installs use
|| trueto tolerate failures (set +e block)
Task 1 [HIGH] ✅ COMPLETED: OpenAI-Compatible Local API Gateway (LiteLLM Proxy)¶
Motivation: Compete with Pi.dev's OpenAI-compatible API. Current state: Ollama/vLLM/SGLang installed but no unified gateway exposing them as OpenAI-compatible endpoints. LiteLLM acts as a proxy/router that unifies all local LLM runtimes behind a single http://localhost:4000/v1 endpoint.
Effort: 2h | Dependencies: None (standalone module)
Files¶
- Create:
src/lib/24-litellm.sh— new module (~60 lines) - Modify:
setup.sh:256,307-308— bump TOTAL_STEPS, add step - Modify:
src/lib/00-core.sh:7— SCRIPT_VERSION → v1.1.0 - Modify:
src/modes/health.sh— add litellm health checks - Modify:
src/lib/19-finalize.sh— add litellm verification
Interfaces¶
- Consumes:
helpers.sh(_curl, log, warn),00-core.sh(step_done, _step_skip) - Produces:
~/.local/bin/litellmbinary,~/.config/litellm/config.yaml, litellm systemd user service
Key Implementation Details¶
#!/usr/bin/env bash
# lib/24-litellm.sh — OpenAI-compatible local API gateway (LiteLLM proxy)
set -euo pipefail
_step_skip step_litellm && return 0
section "OpenAI-compatible API Gateway (LiteLLM)"
# LiteLLM — unified proxy for Ollama/vLLM/SGLang behind OpenAI-compatible /v1
if ! command -v litellm &>/dev/null; then
info "Installing LiteLLM (OpenAI-compatible proxy for local LLMs)..."
pipx install litellm 2>/dev/null && log "LiteLLM installed" || warn "LiteLLM install failed"
fi
# Generate config: auto-detect running backends
mkdir -p ~/.config/litellm
# Detect Ollama models and register them as proxies
OLLAMA_MODELS="qwen3:1.8b"
if command -v ollama &>/dev/null; then
OLLAMA_MODELS=$(ollama list 2>/dev/null | awk 'NR>1{print $1}' | tr '\n' ',' | sed 's/,$//')
[ -z "$OLLAMA_MODELS" ] && OLLAMA_MODELS="qwen3:1.8b"
fi
cat > ~/.config/litellm/config.yaml << YAML
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/LITELLM_DATABASE_URL
model_list:
YAML
# Register Ollama models
for model in $(echo "$OLLAMA_MODELS" | tr ',' ' '); do
cat >> ~/.config/litellm/config.yaml << YAML
- model_name: ollama/${model}
litellm_params:
model: ollama/${model}
api_base: http://localhost:11434
YAML
done
# Register vLLM if available
if command -v vllm &>/dev/null; then
cat >> ~/.config/litellm/config.yaml << YAML
- model_name: vllm-self-hosted
litellm_params:
model: openai/default
api_base: http://localhost:8000/v1
YAML
fi
litellm_version: 1
# Register remote providers (from auth.json) as fallback models
# DeepSeek, OpenCode, xAI — handled by opencode.json provider config
# LiteLLM serves as the LOCAL-only gateway; remote providers stay in opencode.json
cat >> ~/.config/litellm/config.yaml << YAML
router_settings:
routing_strategy: usage-based
enable_pre_call_checks: true
allowed_fails: 3
num_retries: 2
fallbacks:
- ollama/qwen3:1.8b
YAML
log "LiteLLM config written: ~/.config/litellm/config.yaml"
# Install systemd user service for auto-start
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/litellm.service << 'SVC'
[Unit]
Description=LiteLLM — OpenAI-compatible Local API Gateway
After=network.target ollama.service
Wants=ollama.service
[Service]
Type=simple
Environment=LITELLM_MASTER_KEY=sk-local-dev-key-change-me
ExecStart=%h/.local/bin/litellm --config %h/.config/litellm/config.yaml --port 4000
Restart=on-failure
RestartSec=5
[Install]
WantedBy=default.target
SVC
systemctl --user daemon-reload
systemctl --user enable litellm.service
systemctl --user start litellm.service 2>/dev/null || true
log "LiteLLM API: http://localhost:4000/v1 (OpenAI-compatible)"
log "Test: curl http://localhost:4000/v1/models"
_step_done step_litellm
Testing Approach¶
# Unit: bash -n src/lib/24-litellm.sh
# Integration: grep -q "LiteLLM" opencode.json (if we add it as MCP)
# Health: curl -s http://localhost:4000/health/liveliness
# E2E: curl -s http://localhost:4000/v1/models | python3 -c "import json,sys; d=json.load(sys.stdin); assert 'data' in d"
Changes to setup.sh¶
# Line 256: TOTAL_STEPS=24
# After line 307: _run_step step_litellm "LiteLLM API Gateway" "$SCRIPT_DIR/src/lib/24-litellm.sh"
Task 2 [HIGH] ✅ COMPLETED: Hardware Auto-Detection + Optimal Backend Selection¶
Motivation: Compete with Lemonade's auto-detection. Current state: 16-llm.sh detects GPU via lspci | grep -i nvidia and nvidia-smi — NVIDIA only. No NPU detection, no AMD ROCm, no Intel OneAPI, no CPU-optimized backend selection (llama.cpp vs Ollama vs vLLM).
Effort: 1.5h | Dependencies: Task 1 (uses litellm for backend routing)
Files¶
- Modify:
src/lib/16-llm.sh— major rewrite (~52 → ~120 lines) - Modify:
src/lib/24-litellm.sh— use hardware detection results for routing config
Interfaces¶
- Consumes:
helpers.sh,00-core.sh - Produces:
HAS_NVIDIA_GPU,HAS_AMD_GPU,HAS_INTEL_GPU,HAS_NPU,OPTIMAL_BACKENDenv vars,~/.config/litellm/hardware.json
Key Implementation Details¶
# lib/16-llm.sh — revised hardware detection section
detect_hardware() {
# GPU detection
HAS_GPU=false; HAS_NVIDIA_GPU=false; HAS_AMD_GPU=false; HAS_INTEL_GPU=false; HAS_NPU=false
# NVIDIA GPU
if lspci 2>/dev/null | grep -qiE 'nvidia|3D controller.*NVIDIA'; then
HAS_NVIDIA_GPU=true; HAS_GPU=true
log "NVIDIA GPU detected"
# Check CUDA version
nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null && log "NVIDIA driver: $(nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null)"
fi
# AMD GPU (ROCm)
if lspci 2>/dev/null | grep -qiE 'amd/ati|Advanced Micro Devices' || \
[ -d /opt/rocm ] || [ -d /opt/amdgpu ]; then
HAS_AMD_GPU=true; HAS_GPU=true
log "AMD GPU detected (ROCm compatible)"
# Check ROCm version
[ -f /opt/rocm/.info/version ] && log "ROCm: $(cat /opt/rocm/.info/version)"
fi
# Intel GPU (OneAPI/Arc)
if lspci 2>/dev/null | grep -qiE 'Intel.*Graphics|Intel.*Arc'; then
HAS_INTEL_GPU=true; HAS_GPU=true
log "Intel GPU detected (OneAPI compatible)"
fi
# NPU (AMD Ryzen AI, Intel Meteor Lake)
if lspci 2>/dev/null | grep -qiE 'IPU|Neural|NPU|Ryzen AI' || \
[ -e /dev/accel/accel0 ] || [ -e /dev/dri/renderD128 ]; then
HAS_NPU=true
log "NPU detected (AMD Ryzen AI / Intel Meteor Lake)"
fi
# Apple Silicon (M1/M2/M3 — macOS only)
if [ "$(uname -s)" = "Darwin" ] && sysctl -n machdep.cpu.brand_string 2>/dev/null | grep -qi 'Apple'; then
HAS_GPU=true
log "Apple Silicon detected (Metal/MPS acceleration)"
fi
# Determine optimal backend
if $HAS_NVIDIA_GPU; then
OPTIMAL_BACKEND="vllm" # Best performance on NVIDIA
elif $HAS_AMD_GPU; then
OPTIMAL_BACKEND="ollama" # ROCm via Ollama
elif $HAS_INTEL_GPU; then
OPTIMAL_BACKEND="ollama" # OneAPI via Ollama
elif $HAS_NPU; then
OPTIMAL_BACKEND="llama.cpp" # CPU fallback, NPU not yet supported in Ollama
else
OPTIMAL_BACKEND="ollama" # CPU-only works with Ollama
fi
export HAS_NVIDIA_GPU HAS_AMD_GPU HAS_INTEL_GPU HAS_NPU OPTIMAL_BACKEND
# Write hardware profile for LiteLLM routing
mkdir -p ~/.config/litellm
python3 -c "
import json
with open('$HOME/.config/litellm/hardware.json', 'w') as f:
json.dump({
'nvidia_gpu': $HAS_NVIDIA_GPU,
'amd_gpu': $HAS_AMD_GPU,
'intel_gpu': $HAS_INTEL_GPU,
'npu': $HAS_NPU,
'optimal_backend': '$OPTIMAL_BACKEND',
'is_wsl2': '${WSL_DISTRO_NAME:+true}',
'arch': '$(uname -m)'
}, f, indent=2)
"
log "Optimal backend: $OPTIMAL_BACKEND"
}
detect_hardware
# ── Backend-specific installs ─────────────────────────────────────────────
# Ollama — universal (NVIDIA CUDA, AMD ROCm, Intel OneAPI, CPU)
if ! command -v ollama &>/dev/null; then
info "Installing Ollama..."
command -v zstd &>/dev/null || sudo apt-get install -y zstd 2>/dev/null || true
_curl "https://ollama.ai/install.sh" /tmp/ollama-install.sh 2>/dev/null && bash /tmp/ollama-install.sh 2>/dev/null && log "Ollama installed" || warn "Ollama install failed"
rm -f /tmp/ollama-install.sh
if command -v ollama &>/dev/null; then
# Pull optimal model based on hardware
if $HAS_GPU; then
ollama pull qwen3:14b 2>/dev/null & log "Ollama: pulling qwen3:14b (GPU-optimized, background)";
else
ollama pull qwen3:1.8b 2>/dev/null & log "Ollama: pulling qwen3:1.8b (CPU-friendly, background)";
fi
fi
else
log "Ollama already installed"
fi
# vLLM — NVIDIA GPU only, for high-performance inference
if $HAS_NVIDIA_GPU && ! command -v vllm &>/dev/null; then
info "Installing vLLM (NVIDIA GPU inference server)..."
pipx install vllm 2>/dev/null && log "vLLM installed" || warn "vLLM install failed (requires NVIDIA GPU + CUDA)"
fi
# llama.cpp — CPU-optimized, best for NPU/no-GPU scenarios
if ! $HAS_GPU && ! command -v llama.cpp &>/dev/null; then
info "Installing llama.cpp (CPU-optimized inference)..."
if command -v brew &>/dev/null; then
brew install llama.cpp 2>/dev/null && log "llama.cpp installed (brew)" || warn "llama.cpp brew install failed"
else
# Build from source — lightweight
git clone --depth 1 https://github.com/ggerganov/llama.cpp /tmp/llama.cpp-build 2>/dev/null && \
make -C /tmp/llama.cpp-build -j$(nproc) llama-cli 2>/dev/null && \
cp /tmp/llama.cpp-build/llama-cli ~/.local/bin/ && log "llama.cpp built" || warn "llama.cpp build failed"
rm -rf /tmp/llama.cpp-build
fi
fi
# Open WebUI — only install if GPU detected or user explicitly chose LLM
if $HAS_GPU || [ "${INTERACTIVE_DO_LLM:-}" = "true" ]; then
if command -v docker &>/dev/null && ! docker ps --format '{{.Names}}' 2>/dev/null | grep -q 'open-webui'; then
info "Installing Open WebUI (Docker)..."
docker run -d --name open-webui --restart unless-stopped \
-p 3300:8080 -v open-webui:/app/backend/data \
--add-host=host.docker.internal:host-gateway \
ghcr.io/open-webui/open-webui:main 2>/dev/null && log "Open WebUI: http://localhost:3300" || warn "Open WebUI failed"
fi
fi
Changes to 24-litellm.sh¶
The hardware.json written by 16-llm.sh is consumed by LiteLLM config generation in Task 1 — routing strategy picks the optimal backend based on detected hardware.
Testing Approach¶
# Unit: bash -n src/lib/16-llm.sh
# Integration: source 16-llm.sh, verify $OPTIMAL_BACKEND is set
# Health: grep "Optimal backend" setup log
Task 3 [HIGH] ✅ COMPLETED: Lightweight/Headless Mode for CI/CD and Scripts¶
Motivation: Compete with both Lemonade (embeddable mode) and Pi.dev (lightweight CI/CD mode). Current state: only full install mode with all 23 components. No way to get a minimal working AI dev environment without 8 languages and ZSH.
Effort: 1.5h | Dependencies: None (standalone mode)
Files¶
- Create:
src/modes/ci.sh— CI/CD headless mode (~100 lines) - Modify:
setup.sh:44-53,116-118— add--cimode - Modify:
src/lib/00-core.sh:7— version bump
Interfaces¶
- Consumes:
helpers.sh,00-core.sh(step tracking, gate logic) - Produces: minimal
opencode.json, OpenCode CLI + Bun only, no ZSH, no Docker, no GUI
Key Implementation Details¶
#!/usr/bin/env bash
# modes/ci.sh — Lightweight CI/CD headless mode
# Installs: OpenCode CLI + Bun + 3 essential MCPs only
# Use: bash setup.sh --ci
# Non-interactive, zero GUI dependencies
if [ "$MODE" != "ci" ]; then return 0; fi
section "CI/CD Headless Mode — Minimal AI Agent Setup"
log "CI mode: OpenCode CLI + Bun + essential MCPs only"
log "Skipping: ZSH, Docker, Chrome, GUI tools, LLM runtimes, RAG"
TOTAL_STEPS=5
# Only install what CI/CD needs
INTERACTIVE_DO_SYSTEM=false
INTERACTIVE_DO_OPENCODE=true
INTERACTIVE_DO_MCP=true
INTERACTIVE_DO_CHROMADB=false
INTERACTIVE_DO_LLM=false
INTERACTIVE_DO_RAG=false
INTERACTIVE_DO_ZSH=false
# ── Step 1: System essentials ────────────────────────────────────────────
_run_step step_system "System essentials" "$SCRIPT_DIR/src/lib/01-system.sh" || true
# ── Step 2: Node.js (OpenCode depends on Node) ───────────────────────────
INTERACTIVE_DO_NODE=true
_run_step step_node "Node.js" "$SCRIPT_DIR/src/lib/06-node.sh" || true
# ── Step 3: OpenCode CLI + Bun ───────────────────────────────────────────
_run_step step_opencode "OpenCode CLI" "$SCRIPT_DIR/src/lib/11-opencode.sh" || true
# ── Step 4: Essential MCPs only (filesystem, github, context7) ───────────
# Override MCP install to only install CI-essential ones
_info_ci() { info "[CI] $1"; }
CI_MCPS=(
"@modelcontextprotocol/server-filesystem"
"@upstash/context7-mcp"
)
if command -v bun &>/dev/null; then
for pkg in "${CI_MCPS[@]}"; do
_info_ci "Installing $pkg..."
bun add -g "$pkg" 2>/dev/null && log "MCP: $pkg" || warn "MCP: $pkg failed (non-fatal)"
done
fi
# ── Step 5: Minimal opencode.json (CI-optimized) ────────────────────────
_info_ci "Generating minimal opencode.json for CI/CD..."
python3 -c "
import json, os
config = {
'model': os.environ.get('OPENCODE_MODEL', 'deepseek/deepseek-v4-pro'),
'small_model': 'deepseek/deepseek-v4-flash',
'provider': {'deepseek': {}},
'mcp': {},
'plugin': [],
}
# Only add filesystem and context7 MCPs if available
bun_bin = os.path.expanduser('~/.bun/bin')
if os.path.exists(f'{bun_bin}/mcp-server-filesystem'):
config['mcp']['filesystem'] = {
'type': 'local',
'command': [f'{bun_bin}/mcp-server-filesystem'],
'enabled': True
}
if os.path.exists(f'{bun_bin}/c7-mcp-server'):
config['mcp']['context7'] = {
'type': 'local',
'command': [f'{bun_bin}/c7-mcp-server'],
'enabled': True
}
# Write config
out_dir = os.path.expanduser('~/.config/opencode')
os.makedirs(out_dir, exist_ok=True)
with open(f'{out_dir}/opencode.json', 'w') as f:
json.dump(config, f, indent=2)
log('CI opencode.json generated')
" 2>/dev/null || warn "CI opencode.json generation failed"
log "CI/CD headless setup complete"
log "Run: opencode --non-interactive 'check repo for issues'"
echo ""
echo -e " ${GREEN}╔══════════════════════════════════════╗${NC}"
echo -e " ${GREEN}║ CI/CD Headless Mode — Ready ║${NC}"
echo -e " ${GREEN}║ opencode --non-interactive ... ║${NC}"
echo -e " ${GREEN}╚══════════════════════════════════════╝${NC}"
exit 0
Changes to setup.sh¶
# Add --ci flag
--ci) MODE="ci"; shift;;
# Add to CLI help
--ci Headless CI/CD mode: OpenCode CLI + essential MCPs only
# Add to early-exit modes (after line 118)
if [ "$MODE" = "ci" ]; then source "$SCRIPT_DIR/src/modes/ci.sh"; fi
Testing Approach¶
# Unit: bash -n src/modes/ci.sh
# Integration: MODE=ci bash setup.sh --dry-run --ci 2>&1 | grep -q "CI/CD Headless"
# E2E: verify only 5 components installed vs 23 in full mode
Task 4 [MEDIUM]: Multimodal Support (Images, Speech)¶
Motivation: Compete with Lemonade's multimodal capabilities. Current state: text-only LLM runtimes. No support for image generation (stable-diffusion.cpp), speech recognition (whisper.cpp), or image understanding via Ollama vision models.
Effort: 2h | Dependencies: Task 2 (needs hardware detection for optimal backend selection)
Files¶
- Modify:
src/lib/16-llm.sh— add multimodal install section (~40 new lines) - Modify:
src/lib/24-litellm.sh— register multimodal models in config - Modify:
src/modes/health.sh— add multimodal health checks
Interfaces¶
- Consumes:
helpers.sh, hardware detection from16-llm.sh - Produces:
~/.local/bin/whisper-cli,~/.local/bin/sd(stable-diffusion), vision model in Ollama
Key Implementation Details¶
# ── Multimodal: Speech Recognition (whisper.cpp) ─────────────────────────
# Section to add to 16-llm.sh after llama.cpp install
if ! command -v whisper-cli &>/dev/null; then
info "Installing whisper.cpp (speech-to-text)..."
git clone --depth 1 https://github.com/ggerganov/whisper.cpp /tmp/whisper.cpp-build 2>/dev/null && \
(cd /tmp/whisper.cpp-build && bash ./models/download-ggml-model.sh base 2>/dev/null && \
make -j$(nproc) 2>/dev/null && cp main ~/.local/bin/whisper-cli) && \
log "whisper.cpp installed (base model)" || warn "whisper.cpp build failed"
rm -rf /tmp/whisper.cpp-build
else
log "whisper.cpp already installed"
fi
# ── Multimodal: Image Generation (stable-diffusion.cpp) ──────────────────
# CPU-only by default; GPU-accelerated if NVIDIA detected
if ! command -v sd &>/dev/null && ! $HAS_NVIDIA_GPU; then
info "Installing stable-diffusion.cpp (CPU image generation)..."
git clone --depth 1 https://github.com/leejet/stable-diffusion.cpp /tmp/sd.cpp-build 2>/dev/null && \
(cd /tmp/sd.cpp-build && mkdir -p build && cd build && \
cmake .. -DSD_SYSTEM_GGML=OFF 2>/dev/null && \
cmake --build . --config Release -j$(nproc) 2>/dev/null && \
cp bin/sd ~/.local/bin/) && log "stable-diffusion.cpp installed" || warn "sd.cpp build failed"
rm -rf /tmp/sd.cpp-build
fi
# ── Multimodal: Pull vision model for Ollama ─────────────────────────────
if command -v ollama &>/dev/null; then
# Pull a vision-capable model (llava or bakllava)
if ! ollama list 2>/dev/null | grep -q 'llava'; then
info "Pulling vision model (llava:7b)..."
ollama pull llava:7b 2>/dev/null & log "Ollama: pulling llava:7b (vision, background)"
fi
fi
LiteLLM Multimodal Config¶
# In 24-litellm.sh config generation, add:
- model_name: whisper-1
litellm_params:
model: openai/whisper-1
api_base: http://localhost:8081/v1
- model_name: ollama/llava:7b
litellm_params:
model: ollama/llava:7b
api_base: http://localhost:11434
Testing Approach¶
# Unit: bash -n src/lib/16-llm.sh
# Health: check whisper-cli --version, sd --help
# E2E: echo "test" | whisper-cli -m ~/.local/share/whisper/models/ggml-base.bin -f - 2>&1 | grep -q "test"
Task 5 [MEDIUM]: 15+ LLM Providers with Session Switching¶
Motivation: Compete with Pi.dev's 15+ providers. Current state: 6 providers (deepseek, opencode, xai, mimo, moonshot, minimax). Need to add: OpenAI, Anthropic (Claude), Google (Gemini), Mistral, Groq, Together AI, Cohere, Fireworks AI, Together AI, Cerebras, Perplexity, Replicate, HuggingFace.
Effort: 2h | Dependencies: None (standalone module)
Files¶
- Create:
src/lib/25-providers.sh— multi-provider registry + config (~80 lines) - Modify:
src/lib/11-opencode.sh— add provider API key collection - Modify:
src/lib/18-opencode-json.sh— extend _build_providers() with all 15+ providers - Modify:
setup.sh:256,307-308— TOTAL_STEPS bump, add step
Interfaces¶
- Consumes:
helpers.sh,00-core.sh,auth.json - Produces: extended
opencode.jsonwith 15+ providers + fallback chains
Key Implementation Details¶
# lib/25-providers.sh — Multi-provider registry
_step_skip step_providers && return 0
section "Multi-Provider Configuration"
# Provider registry: (short_name api_key_env cli_flag description free_tier)
declare -A PROVIDER_REGISTRY
PROVIDER_REGISTRY=(
[deepseek]="DEEPSEEK_API_KEY|--deepseek-key|DeepSeek V4 Pro (direct)|yes"
[opencode]="OPENCODE_API_KEY|-k|OpenCode Go proxy|yes"
[xai]="XAI_API_KEY|--xai-key|xAI Grok 3|no"
[mimo]="MIMO_API_KEY|--mimo-key|Xiaomi MiMo|yes"
[moonshot]="MOONSHOT_API_KEY|--moonshot-key|Moonshot Kimi K2.6|no"
[minimax]="MINIMAX_API_KEY|--minimax-key|MiniMax M3|no"
[openai]="OPENAI_API_KEY|--openai-key|OpenAI GPT-5|no"
[anthropic]="ANTHROPIC_API_KEY|--anthropic-key|Anthropic Claude 4|no"
[google]="GOOGLE_API_KEY|--google-key|Google Gemini 2.5|yes"
[mistral]="MISTRAL_API_KEY|--mistral-key|Mistral Large 3|no"
[groq]="GROQ_API_KEY|--groq-key|Groq Cloud (fast inference)|yes"
[together]="TOGETHER_API_KEY|--together-key|Together AI|yes"
[cohere]="COHERE_API_KEY|--cohere-key|Cohere Command R+|yes"
[fireworks]="FIREWORKS_API_KEY|--fireworks-key|Fireworks AI|no"
[cerebras]="CEREBRAS_API_KEY|--cerebras-key|Cerebras (fast inference)|no"
[perplexity]="PERPLEXITY_API_KEY|--perplexity-key|Perplexity (online search)|no"
)
# Detect available providers from env vars + auth.json
AVAILABLE_PROVIDERS=""
for provider in "${!PROVIDER_REGISTRY[@]}"; do
IFS='|' read -r env_var _cli_flag _desc _free <<< "${PROVIDER_REGISTRY[$provider]}"
key_val="${!env_var:-}"
# Check auth.json as fallback
if [ -z "$key_val" ] && [ -f "$HOME/.local/share/opencode/auth.json" ]; then
key_val=$(python3 -c "
import json, os
try:
with open('$HOME/.local/share/opencode/auth.json') as f:
auth = json.load(f)
print(auth.get('$provider', {}).get('key', ''))
except: pass
" 2>/dev/null)
fi
if [ -n "$key_val" ]; then
AVAILABLE_PROVIDERS="$AVAILABLE_PROVIDERS $provider"
log "Provider: $provider (API key found)"
fi
done
[ -z "$AVAILABLE_PROVIDERS" ] && warn "No provider API keys found — some features disabled"
export AVAILABLE_PROVIDERS
_step_done step_providers
Extend 18-opencode-json.sh _build_providers()¶
def _build_providers():
"""Build provider config with fallback chains for 15+ providers."""
opts = {"options": {"timeout": 600000, "chunkTimeout": 60000, "setCacheKey": True}}
providers = {}
all_keys = {
"deepseek": os.environ.get("DEEPSEEK_API_KEY") or secrets.get("DEEPSEEK_API_KEY", ""),
"opencode": os.environ.get("OPENCODE_API_KEY") or secrets.get("OPENCODE_API_KEY", ""),
"xai": os.environ.get("XAI_API_KEY") or secrets.get("XAI_API_KEY", ""),
"mimo": os.environ.get("MIMO_API_KEY") or secrets.get("MIMO_API_KEY", ""),
"moonshot": os.environ.get("MOONSHOT_API_KEY") or secrets.get("MOONSHOT_API_KEY", ""),
"minimax": os.environ.get("MINIMAX_API_KEY") or secrets.get("MINIMAX_API_KEY", ""),
"openai": os.environ.get("OPENAI_API_KEY") or secrets.get("OPENAI_API_KEY", ""),
"anthropic": os.environ.get("ANTHROPIC_API_KEY") or secrets.get("ANTHROPIC_API_KEY", ""),
"google": os.environ.get("GOOGLE_API_KEY") or secrets.get("GOOGLE_API_KEY", ""),
"mistral": os.environ.get("MISTRAL_API_KEY") or secrets.get("MISTRAL_API_KEY", ""),
"groq": os.environ.get("GROQ_API_KEY") or secrets.get("GROQ_API_KEY", ""),
"together": os.environ.get("TOGETHER_API_KEY") or secrets.get("TOGETHER_API_KEY", ""),
"cohere": os.environ.get("COHERE_API_KEY") or secrets.get("COHERE_API_KEY", ""),
"fireworks": os.environ.get("FIREWORKS_API_KEY") or secrets.get("FIREWORKS_API_KEY", ""),
"cerebras": os.environ.get("CEREBRAS_API_KEY") or secrets.get("CEREBRAS_API_KEY", ""),
"perplexity": os.environ.get("PERPLEXITY_API_KEY") or secrets.get("PERPLEXITY_API_KEY", ""),
}
provider_models = {
"deepseek": ("deepseek/deepseek-v4-pro", "deepseek/deepseek-v4-flash"),
"openai": ("openai/gpt-5", "openai/gpt-5-mini"),
"anthropic": ("anthropic/claude-sonnet-4-20250514", "anthropic/claude-haiku-4-20250514"),
"google": ("google/gemini-2.5-pro", "google/gemini-2.5-flash"),
"mistral": ("mistral/mistral-large-latest", "mistral/mistral-small-latest"),
"groq": ("groq/llama-4-maverick", "groq/llama-4-scout"),
"together": ("together/meta-llama/Llama-4-Maverick", "together/meta-llama/Llama-4-Scout"),
"xai": ("xai/grok-3", "xai/grok-3-mini"),
"minimax": ("minimax/minimax-m3", "minimax/minimax-m3"),
"mimo": ("mimo/mimo-v2", "mimo/mimo-v2"),
"moonshot": ("moonshot/kimi-k2.6", "moonshot/kimi-k2.6"),
"perplexity": ("perplexity/sonar-pro", "perplexity/sonar"),
}
# Build provider entries with fallback chains
fallback_chain = ["deepseek", "groq", "together", "openai", "minimax"]
for provider, key in all_keys.items():
if key or provider in ("deepseek", "opencode"): # deepseek/opencode always available
providers[provider] = dict(opts)
if provider in provider_models:
providers[provider]["default_model"] = provider_models[provider][0]
providers[provider]["small_model"] = provider_models[provider][1]
# Set fallback (deepseek as universal fallback)
if provider != "deepseek":
providers[provider]["fallback"] = ["deepseek"]
elif provider == "deepseek":
chain = [p for p in fallback_chain if p != "deepseek" and p in providers]
if chain:
providers["deepseek"]["fallback"] = chain[:3]
return providers
Testing Approach¶
# Unit: bash -n src/lib/25-providers.sh src/lib/11-opencode.sh
# Integration: source 25-providers.sh, check $AVAILABLE_PROVIDERS count
# E2E: python3 -c "import json; c=json.load(open('...opencode.json')); assert len(c['provider']) >= 6"
Task 6 [MEDIUM]: Multiple Interaction Modes (TUI/JSON/RPC/SDK)¶
Motivation: Compete with Pi.dev's multiple interaction modes. Current state: only OpenCode CLI interactive mode. Need: TUI (terminal UI), JSON (structured output for scripting), RPC (remote procedure calls for integration), SDK (Python/Node.js wrapper).
Effort: 2h | Dependencies: Task 1 (litellm gateway), Task 5 (providers)
Files¶
- Create:
scripts/opencode-tui.sh— terminal UI wrapper (~50 lines) - Create:
scripts/opencode-json.sh— JSON mode wrapper (~30 lines) - Create:
scripts/opencode-rpc.sh— RPC server wrapper (~40 lines) - Create:
scripts/opencode-sdk.py— Python SDK wrapper (~60 lines) - Modify:
src/lib/11-opencode.sh— add TUI/JSON/RPC install section
Interfaces¶
- Consumes: OpenCode CLI, Bun, LiteLLM (Task 1)
- Produces:
~/.local/bin/oc-tui,~/.local/bin/oc-json,~/.local/bin/oc-rpc, Python SDK
Key Implementation Details¶
# scripts/opencode-tui.sh — Terminal UI for OpenCode
# Uses dialog/whiptail for interactive prompt construction
cat > ~/.local/bin/oc-tui << 'TUI'
#!/usr/bin/env bash
# oc-tui — Terminal UI wrapper for OpenCode
set -euo pipefail
# Check for dialog or whiptail
if command -v dialog &>/dev/null; then
DIALOG="dialog"
elif command -v whiptail &>/dev/null; then
DIALOG="whiptail"
else
echo "Install dialog or whiptail for TUI mode: sudo apt install dialog"
exit 1
fi
# Select model
MODEL=$($DIALOG --title "OpenCode TUI" --menu "Select model:" 15 60 8 \
"deepseek/deepseek-v4-pro" "DeepSeek V4 Pro (recommended)" \
"deepseek/deepseek-v4-flash" "DeepSeek V4 Flash (fast)" \
"openai/gpt-5" "OpenAI GPT-5" \
"anthropic/claude-sonnet" "Anthropic Claude Sonnet 4" \
3>&1 1>&2 2>&3) || exit 0
# Get prompt
PROMPT=$($DIALOG --title "OpenCode TUI" --inputbox "Enter your prompt:" 10 60 3>&1 1>&2 2>&3) || exit 0
# Execute
opencode --model "$MODEL" "$PROMPT"
TUI
chmod +x ~/.local/bin/oc-tui
log "oc-tui installed (~/.local/bin/oc-tui)"
# scripts/opencode-sdk.py — Python SDK for OpenCode automation
# Usage: python3 -c "from opencode_sdk import OpenCode; oc = OpenCode(); print(oc.ask('fix this bug'))"
cat > ~/.local/bin/opencode-sdk.py << 'SDK'
#!/usr/bin/env python3
"""OpenCode SDK — Python wrapper for non-interactive agent use."""
import subprocess, json, os, sys
class OpenCode:
def __init__(self, model=None, api_key=None, workdir=None):
self.model = model or os.environ.get("OPENCODE_MODEL", "deepseek/deepseek-v4-pro")
self.api_key = api_key or os.environ.get("DEEPSEEK_API_KEY", "")
self.workdir = workdir or os.getcwd()
self.cmd = ["opencode", "--model", self.model]
if self.api_key:
os.environ["DEEPSEEK_API_KEY"] = self.api_key
def ask(self, prompt: str, files: list = None) -> dict:
"""Ask OpenCode to perform a task. Returns structured response."""
args = self.cmd + ["--non-interactive", "--json-output", prompt]
if files:
args += ["--files"] + files
result = subprocess.run(args, capture_output=True, text=True, cwd=self.workdir, timeout=300)
try:
return json.loads(result.stdout)
except json.JSONDecodeError:
return {"status": "error", "output": result.stdout, "stderr": result.stderr}
def review(self, filepath: str) -> dict:
"""Review a file for issues."""
return self.ask(f"Review {filepath} for bugs, security issues, and style problems")
def generate(self, spec: str, output_file: str) -> dict:
"""Generate code from spec and write to file."""
return self.ask(f"Generate code for: {spec}. Write the result to {output_file}")
def list_models(self) -> list:
"""List available models from the local API gateway."""
import urllib.request
try:
with urllib.request.urlopen("http://localhost:4000/v1/models") as resp:
data = json.load(resp)
return [m["id"] for m in data.get("data", [])]
except Exception:
return ["deepseek/deepseek-v4-pro"]
# CLI mode
if __name__ == "__main__":
if len(sys.argv) < 2:
print("Usage: opencode-sdk.py '<prompt>'")
sys.exit(1)
oc = OpenCode()
result = oc.ask(sys.argv[1])
print(json.dumps(result, indent=2))
SDK
chmod +x ~/.local/bin/opencode-sdk.py
log "Python SDK installed (~/.local/bin/opencode-sdk.py)"
Testing Approach¶
# Unit: bash -n scripts/opencode-tui.sh scripts/opencode-json.sh scripts/opencode-rpc.sh
# Integration: python3 -c "from opencode_sdk import OpenCode; oc = OpenCode(); assert oc.model"
Task 7 [LOW]: Desktop UI / Web Interface for Model Management¶
Motivation: Compete with Lemonade's desktop UI. Current state: Open WebUI provides chat interface, but no model management dashboard. Need: model download/delete/configure UI.
Effort: 3h | Dependencies: Task 1 (litellm), Task 2 (hardware detection)
Files¶
- Create:
src/lib/26-model-ui.sh— model management web UI module (~80 lines) - Modify:
setup.sh— add step
Interfaces¶
- Consumes:
helpers.sh, LiteLLM, Ollama - Produces: Open WebUI with admin features enabled, model management API
Key Implementation Details¶
Rather than building a new UI from scratch, this task extends the existing Open WebUI deployment with admin features and model management pipelines:
# Enable Open WebUI admin features for model management
# Configure model download presets based on hardware detection
# Add model download queue for background pulls
# Expose Open WebUI on all interfaces (0.0.0.0) for LAN access
Testing Approach¶
# Manual: open http://localhost:3300/admin/models in browser
# Health: curl -s http://localhost:3300/api/models | python3 -c "import json,sys; print(len(json.load(sys.stdin)['data']))"
Task 8 [LOW]: ONNX Runtime Support¶
Motivation: Compete with Lemonade's ONNX runtime support. Current state: Ollama (llama.cpp backend), vLLM, SGLang — no ONNX. ONNX enables cross-platform model portability and hardware vendor neutrality.
Effort: 2h | Dependencies: Task 2 (hardware detection)
Files¶
- Modify:
src/lib/16-llm.sh— add ONNX runtime install section (~30 lines)
Interfaces¶
- Consumes: hardware detection
- Produces:
~/.local/bin/onnxruntime_perf_test, ONNX model cache
Key Implementation Details¶
# ── ONNX Runtime ─────────────────────────────────────────────────────────
if ! python3 -c "import onnxruntime" 2>/dev/null; then
info "Installing ONNX Runtime (cross-platform inference)..."
pipx install onnxruntime 2>/dev/null && log "ONNX Runtime installed" || \
pip install --user onnxruntime 2>/dev/null && log "ONNX Runtime installed (pip)" || \
warn "ONNX Runtime install failed"
fi
# Pull a small ONNX model as proof-of-concept
MODEL_CACHE="$HOME/.cache/opencode-setup/models"
mkdir -p "$MODEL_CACHE"
if [ ! -f "$MODEL_CACHE/all-MiniLM-L6-v2.onnx" ]; then
info "Downloading ONNX embedding model (MiniLM-L6)..."
_curl "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" \
"$MODEL_CACHE/all-MiniLM-L6-v2.onnx" 2>/dev/null && log "ONNX model cached" || warn "ONNX model download failed"
fi
Testing Approach¶
# Unit: python3 -c "import onnxruntime; print(onnxruntime.__version__)"
# Health: check onnxruntime via python import
Task 9 [LOW]: Embeddable Lightweight Mode for CI/CD Pipelines¶
Motivation: Extend Task 3 (CI/CD headless mode) with Docker-less, pip-less, truly embedded mode — single binary deployment. Compete with Lemonade's embeddable mode.
Effort: 2h | Dependencies: Task 3 (CI mode)
Files¶
- Modify:
src/modes/ci.sh— add--embeddedsub-flag (~40 lines) - Create:
scripts/build-embedded.sh— static build script (~50 lines)
Interfaces¶
- Produces: single
opencode-embeddedbinary using Bun compile
Key Implementation Details¶
# scripts/build-embedded.sh
# Uses bun build --compile to create a single binary
# Includes: OpenCode CLI + essential MCPs + bun runtime
# Output: ~200MB standalone binary, no dependencies
Testing Approach¶
# Build: bash scripts/build-embedded.sh
# Test: ./opencode-embedded --version
# Test: ./opencode-embedded --non-interactive "echo hello"
Dependency Graph¶
Task 2 (hardware detection)
├─→ Task 1 (litellm gateway) ←─ depends on hardware.json
├─→ Task 4 (multimodal) ←─ needs hardware for model selection
└─→ Task 8 (ONNX) ←─ needs hardware for backend choice
Task 1 (litellm) + Task 5 (providers)
└─→ Task 6 (TUI/JSON/RPC) ←─ needs gateway + multiple providers
Task 3 (CI mode)
└─→ Task 9 (embedded) ←─ extends CI mode
Task 7 (Desktop UI) ←─ independent
HIGH priority tasks (1, 2, 3) are independent and can be done in parallel.
Execution Order (Recommended)¶
Session 1 (HIGH): Task 2 → Task 1 → Task 3 (hardware → gateway → CI mode)
Session 2 (MEDIUM): Task 4 + Task 5 → Task 6 (multimodal + providers → interaction modes)
Session 3 (LOW): Task 8 → Task 7 → Task 9 (ONNX → desktop UI → embedded)
Verification Checklist¶
After HIGH tasks (1-3): - [x] bash -n passes on all src/lib/*.sh src/modes/*.sh setup.sh dev.sh - [x] ShellCheck passes on all .sh files - [x] bash setup.sh --health shows 65+ passed (up from 60+) - [x] curl -s http://localhost:4000/health/liveliness returns OK - [x] opencode --version works in CI mode - [x] All CLI flags documented in setup.sh --help - [x] CHANGELOG.md updated with v1.1.0 entry - [x] README.md updated with new capabilities - [x] AGENTS.md version table updated
After ALL tasks: - [ ] bash tests/run_tests.sh passes all tests (updated assertions) - [ ] python3 -c "import json; json.load(open('~/.config/opencode/opencode.json'))" succeeds - [ ] python3 -c "from opencode_sdk import OpenCode" succeeds - [ ] ollama list | grep llava shows vision model - [ ] python3 -c "import onnxruntime" succeeds