OpenCode on GitHub with Kubernetes

OpenCode on GitHub with Kubernetes

In this article, we’ll learn how to use OpenCode to automatically make changes to code using GitHub issues and pull requests. You’ll see how to use various AI models, including custom ones running on a Kubernetes cluster. The OpenCode application installed on GitHub will automatically analyze our issues and pull requests. It will then communicate with the AI model on Anthropic and with our Qwen3 Coder model running on Kubernetes via vLLM. GitHub Actions will perform all necessary tasks.

If you’re interested in AI solutions related to code generation, you might want to check out a few other articles on my blog. The first is the post “Continuous Development with Claude Code on GitHub”, which describes a use case similar to this article but using Claude Code. In the following article, you’ll learn how to create a template containing Claude-compatible skills for building Spring Boot applications. And if you’re interested in running Java applications on Kubernetes, you can find a lot more hands-on examples in my book “Hands-On Java with Kubernetes”.

Source Code

Feel free to use my source code if you’d like to try the exercise yourself. To do that, you must clone my sample GitHub repository. Then follow my instructions.

Actually, this isn’t the only repository where I use OpenCode. Many of them also run Renovate, which updates the list of dependencies in pom.xml. Each time, it opens a pull request and triggers the defined workflows. You can take a look, for example, at the list of issues and pull requests in repositories such as sample-spring-elasticsearch or sample-spring-microservices-advanced.

Motivation

The main goal of this exercise is to create an automated development process in a GitHub repository using AI and OpenCode. This way, to change the source code, we’ll just describe the task in an issue and then merge the pull request into the main branch. After installing the OpenCode application on GitHub, an agent will handle both development and review. Individual GitHub Actions workflows are responsible for triggering it. Additionally, I want to show you how to use various AI models, not just the most obvious ones provided, e.g., those from Anthropic.

Run OpenCode on GitHub

Install the OpenCode GitHub App

First, install the OpenCode Agent app in your GitHub account. Then click the Configure button.

In the configuration form, choose the repositories included in the process. Alternatively, you can also enable OpenCode for all repositories.

opencode-github-installation

Create GitHub Workflow

The decision is yours. You can also install OpenCode locally and then run the command opencode github install in the relevant repository. OpenCode will ask you to select a target model. However, this does not change the fact that, ultimately, a GitHub Actions configuration file will be created. So you can simply create a file like the one below right away in the .github/workflows directory in your repository.

The first workflow listens for issue comments and pull request review comments. When a comment body contains /oc or /opencode. OpenCode is invoked to act on the instruction. At this stage, all of our workflows use either the Claude Opus or Sonnet model (google-vertex-anthropic/claude-sonnet-4-6@default), available through Google Vertex AI. So you need a GCP project with Vertex AI enabled and a service account with the appropriate IAM roles.

name: opencode

on:
  issue_comment:
    types: [created]
  pull_request_review_comment:
    types: [created]

jobs:
  opencode:
    if: |
      contains(github.event.comment.body, ' /oc') ||
      startsWith(github.event.comment.body, '/oc') ||
      contains(github.event.comment.body, ' /opencode') ||
      startsWith(github.event.comment.body, '/opencode')
    runs-on: ubuntu-latest
    permissions:
      id-token: write
      contents: read
      pull-requests: read
      issues: read
    steps:
      - name: Checkout repository
        uses: actions/checkout@v7
        with:
          persist-credentials: false

      - name: Create Google credentials
        env:
          GCP_SERVICE_ACCOUNT_JSON: ${{ secrets.GCP_SERVICE_ACCOUNT_JSON }}
        run: |
          printf '%s' "$GCP_SERVICE_ACCOUNT_JSON" > "$RUNNER_TEMP/gcp-credentials.json"

      - name: Run opencode
        uses: anomalyco/opencode/github@latest
        env:
          GOOGLE_VERTEX_PROJECT: ${{ secrets.GOOGLE_VERTEX_PROJECT }}
          GOOGLE_VERTEX_LOCATION: ${{ secrets.GOOGLE_VERTEX_LOCATION }}
          GOOGLE_APPLICATION_CREDENTIALS: ${{ runner.temp }}/gcp-credentials.json
        with:
          model: google-vertex-anthropic/claude-sonnet-4-6@default
YAML

The case we are considering is not the simplest. The workflow shown above uses several GitHub secrets. Access to the model via Google Vertex AI is granted using GCP credentials, which we must enter into the workflow. You’ll need to read the contents of the JSON file with your GCP account credentials. To do that, run the following command. Then, copy that content and save it as the GCP_SERVICE_ACCOUNT_JSON secret.

cat .config/gcloud/application_default_credentials.json
ShellSession

Then you must create secrets for GitHub Actions in your repository. Go to the Settings / Secrets and variables / Actions section. Those secrets should contain Google’s project name, location, and the contents of the JSON file downloaded in the previous step.

Handle GitHub Issues with OpenCode Agent

That’s basically all we had to do to see how the OpenCode GitHub app works. Now we can create a new issue on GitHub. As you can see below, it’s a very simple task. After we create and describe our issue, we need to start our comment with /oc or /opencode. The agent then executes the task. After a moment, you should see the result as a pull request created by that agent.

opencode-github-issue

The OpenCode agent is triggered by GitHub Actions, which triggers on an issue comment.

You can easily view details about the agent’s execution and its interaction with the model. Just go to the selected GitHub workflow and expand the logs for the relevant step.

Here’s what our pull request looks like. As you can see, OpenCode Agent created a new branch, submitted a pull request, and described the code changes. Next, it initiated the review process. But more on that later in this article.

opencode-github-pull-request

Running a Custom AI Model on Kubernetes

Install Operators for GPU on OpenShift

In the previous section, you learned how to use the OpenCode app in your GitHub repository. The OpenCode app automatically analyzed our issue, created a pull request with the changes, which we then merged into the main branch. So far, we’ve used the Anthropic model managed by Google Vertex AI. In this section, you’ll see how to switch to a model running and deployed on Kubernetes. More specifically, we’ll deploy our OpenCode app’s model on OpenShift using vLLM.

To do that, we’ll start by installing and configuring two Kubernetes operators. These operators are the NVIDIA GPU Operator and the Node Feature Discovery Operator. In OpenShift, installing an operator is a one-click process, followed by a few moments of waiting. If you see what I see, you can move on.

opencode-github-openshift-operators

The Node Feature Discovery Operator will be installed in the openshift-nfd namespace. Then go to the NodeFeatureDiscovery tab and create an object with the same name by clicking the button below. You don’t need to change anything. Just accept the default field values.

Then do a very similar thing with the NVIDIA GPU Operator. You will find it in the nvidia-gpu-operator namespace. Create the ClusterPolicy object by clicking the button visible below. As before, don’t change anything. Leave all field values at their defaults. If the object’s status is “ready” after it’s created, we can basically call it a success. However, still not a complete success 🙂

Provision node with GPU

My OpenShift cluster is running on AWS. It has a default code pool, and generally everything works fine on it. But to run even a slightly larger AI model, we need a more powerful machine than a standard node. First, we need a machine with a GPU, which doesn’t make sense to use as a standard OpenShift node. To create a new, custom node in OpenShift, we need to define what’s called a MachineSet. Poniżej komenda, za pomocą której tworzę MachineSet, wykorzystując maszyny g6.12xlarge na AWS.

rosa create machinepool \
  --cluster=piomin-aws-main \
  --name=ai-small \
  --instance-type=g6.12xlarge \
  --replicas=1
ShellSession

Why the g6.12xlarge instance? The model we’ll be running today is Qwen3 Coder Next. The g6.12xlarge instance is essentially the smallest AWS instance on which this model can run effectively. It has 4 GPUs and 96 GB of VRAM. Later on, you’ll see how to run this model on vLLM so that it uses the 4 available GPUs. Once the node is up and running, use this command to check whether Kubernetes detected the GPU there.

Run Model using vLLM

Now that our GPU node is up and running, we can move on to running an AI model on it. You can find the model we’ll be using in the Hugging Face repository here: RedHatAI/Qwen3-Coder-Next-NVFP4. This model uses FP4 quantization for the weights and activations of Qwen/Qwen3-Coder-Next. The optimization reduces the number of bits per parameter from 16 to 4, cutting the model’s disk footprint and GPU memory requirements by approximately 75%.

The YAML defines a Kubernetes Deployment that runs Qwen3-Coder-Next with vLLM in the ai namespace. The workload requests four NVIDIA GPUs, providing the compute resources required for inference. It uses Red Hat’s latest stable vLLM CUDA image for RHEL 9, which provides the vLLM runtime and CUDA environment for GPU inference. The deployment loads the RedHatAI/Qwen3-Coder-Next-NVFP4 model and distributes it across all four GPUs using tensor parallelism. Finally, it exposes vLLM’s OpenAI-compatible API on port 8000. and enables tool calling with automatic tool selection.

apiVersion: apps/v1
kind: Deployment
metadata:
  namespace: ai
  name: qwen3-coder
  annotations: {}
spec:
  selector:
    matchLabels:
      app: qwen3-coder
  replicas: 1
  template:
    metadata:
      labels:
        app: qwen3-coder
    spec:
      containers:
        - resources:
            limits:
              cpu: '16'
              memory: 50Gi
              nvidia.com/gpu: '4'
            requests:
              cpu: '1'
              memory: 20Gi
              nvidia.com/gpu: '4'
          name: vllm
          image: registry.redhat.io/rhaii/vllm-cuda-rhel9:3.4.5
          command:
            - python
            - '-m'
            - vllm.entrypoints.openai.api_server
          args:
            - '--port=8000'
            - '--model=RedHatAI/Qwen3-Coder-Next-NVFP4'
            - '--served-model-name=qwen3_coder'
            - '--tensor-parallel-size=4'
            - '--tool-call-parser=qwen3_coder'
            - '--enable-auto-tool-choice'
            - '--enforce-eager'
          ports:
            - containerPort: 8000
              protocol: TCP
          env:
            - name: HF_HUB_OFFLINE
              value: '0'
            - name: HUGGING_FACE_HUB_TOKEN
              value: <YOUR_HUGGING_FACE_TOKEN>
YAML

The base image names or runtime parameters may change. You can usually find the most up-to-date version in my repository, in the ai directory.

To expose the model outside of the OpenShift cluster, we apply the following to the YAMLs.

apiVersion: v1
kind: Service
metadata:
  name: qwen3-coder
  namespace: ai
spec:
  selector:
    app: qwen3-coder
  ports:
    - protocol: TCP
      port: 8000
      targetPort: 8000
---
kind: Route
apiVersion: route.openshift.io/v1
metadata:
  name: qwen3-coder
  namespace: ai
spec:
  to:
    kind: Service
    name: qwen3-coder
    weight: 100
  port:
    targetPort: 8000
YAML

Now we need to be patient. It may take a few minutes to bring up the model. Eventually, however, you should see logs like the ones below. This means we’ve just launched Qwen3 Coder on Kubernetes!

opencode-github-qwen3-coder-logs

Our model is accessible from outside the cluster at the address Route qwen3-coder. To check that URL on your cluster, run the following command:

oc get route qwen3-coder -n ai
ShellSession

Integrate Custom Model with OpenCode on GitHub

Now that we’ve deployed the model externally, we can integrate it with OpenCode. First, define a new, custom OpenAI-compatible model in the GitHub repository. To do this, we’ll create a configuration file opencode.json as shown below. It must contain the provider name, the URL of the inference server, and the model name. The model name must match the name we set in its Deployment under the --served-model-name parameter. The provider name, on the other hand, can be anything, but we’ll use it later when calling the model from the GitHub application.

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "vllm": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Qwen3 Coder Next",
      "options": {
        "baseURL": "http://qwen3-coder-ai.apps.piomin-aws-main.5fuc.p1.openshiftapps.com/v1"
      },
      "models": {
        "qwen3_coder": {
          "name": "Qwen3 Coder Next"
        }
      }
    }
  }
}
JSON

Now, we can go back to defining our GitHub workflow. It will be simpler than before for the Google Vertex AI-based model, because our endpoint does not require authorization. In the Run OpenCode section, set the provider name and the model name.

name: opencode

on:
  issue_comment:
    types: [created]
  pull_request_review_comment:
    types: [created]

jobs:
  opencode:
    if: |
      contains(github.event.comment.body, ' /oc') ||
      startsWith(github.event.comment.body, '/oc') ||
      contains(github.event.comment.body, ' /opencode') ||
      startsWith(github.event.comment.body, '/opencode')
    runs-on: ubuntu-latest
    permissions:
      id-token: write
      contents: read
      pull-requests: read
      issues: read
    steps:
      - name: Checkout repository
        uses: actions/checkout@v7
        with:
          persist-credentials: false

      - name: Run opencode
        uses: anomalyco/opencode/github@latest
        with:
          model: vllm/qwen3_coder
YAML

Next, if you add a comment starting with /opencode to any issue in your GitHub repository, the AI agent will communicate with the custom model. In this case, you should see a communication trace in the qwen3-coder Pod logs.

OpenCode for Pull Requests Review

Introduction

In this section, we’ll analyze a slightly more advanced scenario involving code changes. In addition to implementing the changes based on the issue description, I’d like OpenCode to review the pull request that was created. This could be a pull request created by a developer or by a bot, such as Renovate or OpenCode Agent. For example, we can use a different model to review a pull request than to implement the task itself.

To do this, we’ll create additional GitHub Actions workflows. But first, I’d like to flag an OpenCode issue that will affect how we implement these workflows. You can read more about it in this OpenCode repository (as well as in a few others): https://github.com/anomalyco/opencode/issues/17224.

In short: if a pull request was created by a bot rather than a user, reviewing it using the OpenCode agent will simply result in a permission error. Perhaps OpenCode will resolve this issue at some point. For now, I’ll suggest a workaround. Our workflow will check the username of the user who created the pull request. It will then handle the pull request differently depending on whether a bot or a standard user created it.

Create and Run GitHub Review Workflows

The GitHub workflow visible below runs the OpenCode review for a standard user. For a bot user, it calls another workflow using the GitHub CLI.

name: Trigger OpenCode review

on:
  pull_request:
    types: [opened]

jobs:
  trigger:
    if: endsWith(github.event.pull_request.user.login, '[bot]')
    runs-on: ubuntu-latest
    permissions:
      actions: write
    steps:
      - name: Trigger OpenCode Review [Bot]
        env:
          GH_TOKEN: ${{ github.token }}
          REPO: ${{ github.repository }}
          PR_NUMBER: ${{ github.event.pull_request.number }}
        run: |
          gh workflow run opencode-review.yml \
            --repo "$REPO" \
            --ref "${{ github.event.repository.default_branch }}" \
            -f pr_number="$PR_NUMBER"

  human-pr-review:
    if: "!endsWith(github.event.pull_request.user.login, '[bot]')"
    runs-on: ubuntu-latest
    permissions:
      id-token: write
      contents: read
      pull-requests: write
      issues: write
    steps:
      - uses: actions/checkout@v7
        with:
          persist-credentials: false

      - name: Create Google credentials
        env:
          GCP_SERVICE_ACCOUNT_JSON: ${{ secrets.GCP_SERVICE_ACCOUNT_JSON }}
        run: |
          set -euo pipefail
          printf '%s' "$GCP_SERVICE_ACCOUNT_JSON" \
            > "$RUNNER_TEMP/gcp-credentials.json"

      - uses: anomalyco/opencode/github@latest
        env:
          GOOGLE_VERTEX_PROJECT: ${{ secrets.GOOGLE_VERTEX_PROJECT }}
          GOOGLE_VERTEX_LOCATION: ${{ secrets.GOOGLE_VERTEX_LOCATION }}
          GOOGLE_APPLICATION_CREDENTIALS: ${{ runner.temp }}/gcp-credentials.json
          GITHUB_TOKEN: ${{ github.token }}
        with:
          model: google-vertex-anthropic/claude-opus-4-6@default
          use_github_token: true
          prompt: |
            Review this pull request for:
            - potential bugs
            - correctness issues
            - security problems
            - code quality issues
            - problematic edge cases
            - unnecessary complexity
            
            Focus on actionable findings that are relevant to the
            changes in this pull request.
            
            Post the review findings.
opencode-review-trigger.yml

Here, in turn, you can see the workflow responsible for reviewing a pull request created by a bot. It was triggered by opencode-review-trigger.yml. As you can see, this workflow uses the pull request number. It checks out the pull request branch and then reviews the changes using the OpenCode agent, which communicates with the vllm/qwen3_coder model.

name: OpenCode review

on:
  workflow_dispatch:
    inputs:
      pr_number:
        description: "Pull request number"
        required: true
        type: string

jobs:
  review:
    runs-on: ubuntu-latest

    permissions:
      id-token: write
      contents: read
      pull-requests: write
      issues: write

    steps:
      - name: Get PR information
        id: pr
        env:
          GH_TOKEN: ${{ github.token }}
          PR_NUMBER: ${{ inputs.pr_number }}
          REPO: ${{ github.repository }}
        run: |
          set -euo pipefail
    
          echo "number=$PR_NUMBER" >> "$GITHUB_OUTPUT"
    
          gh pr view "$PR_NUMBER" \
            --repo "$REPO" \
            --json url,title,headRefName,baseRefName \
            > "$RUNNER_TEMP/pr.json"
    
          cat "$RUNNER_TEMP/pr.json"

      - name: Checkout PR
        uses: actions/checkout@v7
        with:
          ref: refs/pull/${{ inputs.pr_number }}/head
          persist-credentials: false

      - name: Run OpenCode
        uses: anomalyco/opencode/github@latest
        env:
          GITHUB_TOKEN: ${{ github.token }}
        with:
          model: vllm/qwen3_coder
          use_github_token: true
          prompt: |
            Review pull request #${{ inputs.pr_number }}.
    
            Repository: ${{ github.repository }}
    
            Review the checked-out pull request for:
            - potential bugs
            - correctness issues
            - security problems
            - code quality issues
            - problematic edge cases
            - unnecessary complexity
    
            Focus on actionable findings that are relevant to the
            changes in this pull request.
    
            Post the review findings to pull request #${{ inputs.pr_number }}.
opencode-review.yml

In general, I’m switching many of my repositories to a maintenance model using the OpenCode app on GitHub. That’s why you’ll easily find examples of issues and pull requests with automatic reviews performed by the OpenCode agent. You can see an example pull request below. It’s located in the https://github.com/piomin/sample-graphql-microservices repository. In such a pull request, CircleCI pipelines are triggered to perform the build and scan the code in SonarCloud. Additionally, a GitHub action is triggered that reviews the changes made by the agent itself 🙂

Here you can see the result of such a review.

Conclusion

OpenCode, GitHub Actions and Kubernetes show how AI can become a natural part of the development workflow — from writing code to reviewing pull requests. Running custom models on Kubernetes gives us even more flexibility: we keep control of our data, models, and infrastructure, while scaling AI workloads as needed. In practice, this turns the AI model into another internal service we can plug into our development tools and adapt to our needs.

Leave a Reply