OpenCode on GitHub with Kubernetes
In this article, we’ll learn how to use OpenCode to automatically make changes to code using GitHub issues and pull requests. You’ll see how to use various AI models, including custom ones running on a Kubernetes cluster. The OpenCode application installed on GitHub will automatically analyze our issues and pull requests. It will then communicate with the AI model on Anthropic and with our Qwen3 Coder model running on Kubernetes via vLLM. GitHub Actions will perform all necessary tasks.
If you’re interested in AI solutions related to code generation, you might want to check out a few other articles on my blog. The first is the post “Continuous Development with Claude Code on GitHub”, which describes a use case similar to this article but using Claude Code. In the following article, you’ll learn how to create a template containing Claude-compatible skills for building Spring Boot applications. And if you’re interested in running Java applications on Kubernetes, you can find a lot more hands-on examples in my book “Hands-On Java with Kubernetes”.
Source Code
Feel free to use my source code if you’d like to try the exercise yourself. To do that, you must clone my sample GitHub repository. Then follow my instructions.
Actually, this isn’t the only repository where I use OpenCode. Many of them also run Renovate, which updates the list of dependencies in pom.xml. Each time, it opens a pull request and triggers the defined workflows. You can take a look, for example, at the list of issues and pull requests in repositories such as sample-spring-elasticsearch or sample-spring-microservices-advanced.
Motivation
The main goal of this exercise is to create an automated development process in a GitHub repository using AI and OpenCode. This way, to change the source code, we’ll just describe the task in an issue and then merge the pull request into the main branch. After installing the OpenCode application on GitHub, an agent will handle both development and review. Individual GitHub Actions workflows are responsible for triggering it. Additionally, I want to show you how to use various AI models, not just the most obvious ones provided, e.g., those from Anthropic.
Run OpenCode on GitHub
Install the OpenCode GitHub App
First, install the OpenCode Agent app in your GitHub account. Then click the Configure button.

In the configuration form, choose the repositories included in the process. Alternatively, you can also enable OpenCode for all repositories.

Create GitHub Workflow
The decision is yours. You can also install OpenCode locally and then run the command opencode github install in the relevant repository. OpenCode will ask you to select a target model. However, this does not change the fact that, ultimately, a GitHub Actions configuration file will be created. So you can simply create a file like the one below right away in the .github/workflows directory in your repository.
The first workflow listens for issue comments and pull request review comments. When a comment body contains /oc or /opencode. OpenCode is invoked to act on the instruction. At this stage, all of our workflows use either the Claude Opus or Sonnet model (google-vertex-anthropic/claude-sonnet-4-6@default), available through Google Vertex AI. So you need a GCP project with Vertex AI enabled and a service account with the appropriate IAM roles.
name: opencode
on:
issue_comment:
types: [created]
pull_request_review_comment:
types: [created]
jobs:
opencode:
if: |
contains(github.event.comment.body, ' /oc') ||
startsWith(github.event.comment.body, '/oc') ||
contains(github.event.comment.body, ' /opencode') ||
startsWith(github.event.comment.body, '/opencode')
runs-on: ubuntu-latest
permissions:
id-token: write
contents: read
pull-requests: read
issues: read
steps:
- name: Checkout repository
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Create Google credentials
env:
GCP_SERVICE_ACCOUNT_JSON: ${{ secrets.GCP_SERVICE_ACCOUNT_JSON }}
run: |
printf '%s' "$GCP_SERVICE_ACCOUNT_JSON" > "$RUNNER_TEMP/gcp-credentials.json"
- name: Run opencode
uses: anomalyco/opencode/github@latest
env:
GOOGLE_VERTEX_PROJECT: ${{ secrets.GOOGLE_VERTEX_PROJECT }}
GOOGLE_VERTEX_LOCATION: ${{ secrets.GOOGLE_VERTEX_LOCATION }}
GOOGLE_APPLICATION_CREDENTIALS: ${{ runner.temp }}/gcp-credentials.json
with:
model: google-vertex-anthropic/claude-sonnet-4-6@defaultYAMLThe case we are considering is not the simplest. The workflow shown above uses several GitHub secrets. Access to the model via Google Vertex AI is granted using GCP credentials, which we must enter into the workflow. You’ll need to read the contents of the JSON file with your GCP account credentials. To do that, run the following command. Then, copy that content and save it as the GCP_SERVICE_ACCOUNT_JSON secret.
cat .config/gcloud/application_default_credentials.jsonShellSessionThen you must create secrets for GitHub Actions in your repository. Go to the Settings / Secrets and variables / Actions section. Those secrets should contain Google’s project name, location, and the contents of the JSON file downloaded in the previous step.

Handle GitHub Issues with OpenCode Agent
That’s basically all we had to do to see how the OpenCode GitHub app works. Now we can create a new issue on GitHub. As you can see below, it’s a very simple task. After we create and describe our issue, we need to start our comment with /oc or /opencode. The agent then executes the task. After a moment, you should see the result as a pull request created by that agent.

The OpenCode agent is triggered by GitHub Actions, which triggers on an issue comment.

You can easily view details about the agent’s execution and its interaction with the model. Just go to the selected GitHub workflow and expand the logs for the relevant step.

Here’s what our pull request looks like. As you can see, OpenCode Agent created a new branch, submitted a pull request, and described the code changes. Next, it initiated the review process. But more on that later in this article.

Running a Custom AI Model on Kubernetes
Install Operators for GPU on OpenShift
In the previous section, you learned how to use the OpenCode app in your GitHub repository. The OpenCode app automatically analyzed our issue, created a pull request with the changes, which we then merged into the main branch. So far, we’ve used the Anthropic model managed by Google Vertex AI. In this section, you’ll see how to switch to a model running and deployed on Kubernetes. More specifically, we’ll deploy our OpenCode app’s model on OpenShift using vLLM.
To do that, we’ll start by installing and configuring two Kubernetes operators. These operators are the NVIDIA GPU Operator and the Node Feature Discovery Operator. In OpenShift, installing an operator is a one-click process, followed by a few moments of waiting. If you see what I see, you can move on.

The Node Feature Discovery Operator will be installed in the openshift-nfd namespace. Then go to the NodeFeatureDiscovery tab and create an object with the same name by clicking the button below. You don’t need to change anything. Just accept the default field values.

Then do a very similar thing with the NVIDIA GPU Operator. You will find it in the nvidia-gpu-operator namespace. Create the ClusterPolicy object by clicking the button visible below. As before, don’t change anything. Leave all field values at their defaults. If the object’s status is “ready” after it’s created, we can basically call it a success. However, still not a complete success 🙂

Provision node with GPU
My OpenShift cluster is running on AWS. It has a default code pool, and generally everything works fine on it. But to run even a slightly larger AI model, we need a more powerful machine than a standard node. First, we need a machine with a GPU, which doesn’t make sense to use as a standard OpenShift node. To create a new, custom node in OpenShift, we need to define what’s called a MachineSet. Poniżej komenda, za pomocą której tworzę MachineSet, wykorzystując maszyny g6.12xlarge na AWS.
rosa create machinepool \
--cluster=piomin-aws-main \
--name=ai-small \
--instance-type=g6.12xlarge \
--replicas=1ShellSessionWhy the g6.12xlarge instance? The model we’ll be running today is Qwen3 Coder Next. The g6.12xlarge instance is essentially the smallest AWS instance on which this model can run effectively. It has 4 GPUs and 96 GB of VRAM. Later on, you’ll see how to run this model on vLLM so that it uses the 4 available GPUs. Once the node is up and running, use this command to check whether Kubernetes detected the GPU there.

Run Model using vLLM
Now that our GPU node is up and running, we can move on to running an AI model on it. You can find the model we’ll be using in the Hugging Face repository here: RedHatAI/Qwen3-Coder-Next-NVFP4. This model uses FP4 quantization for the weights and activations of Qwen/Qwen3-Coder-Next. The optimization reduces the number of bits per parameter from 16 to 4, cutting the model’s disk footprint and GPU memory requirements by approximately 75%.
The YAML defines a Kubernetes Deployment that runs Qwen3-Coder-Next with vLLM in the ai namespace. The workload requests four NVIDIA GPUs, providing the compute resources required for inference. It uses Red Hat’s latest stable vLLM CUDA image for RHEL 9, which provides the vLLM runtime and CUDA environment for GPU inference. The deployment loads the RedHatAI/Qwen3-Coder-Next-NVFP4 model and distributes it across all four GPUs using tensor parallelism. Finally, it exposes vLLM’s OpenAI-compatible API on port 8000. and enables tool calling with automatic tool selection.
apiVersion: apps/v1
kind: Deployment
metadata:
namespace: ai
name: qwen3-coder
annotations: {}
spec:
selector:
matchLabels:
app: qwen3-coder
replicas: 1
template:
metadata:
labels:
app: qwen3-coder
spec:
containers:
- resources:
limits:
cpu: '16'
memory: 50Gi
nvidia.com/gpu: '4'
requests:
cpu: '1'
memory: 20Gi
nvidia.com/gpu: '4'
name: vllm
image: registry.redhat.io/rhaii/vllm-cuda-rhel9:3.4.5
command:
- python
- '-m'
- vllm.entrypoints.openai.api_server
args:
- '--port=8000'
- '--model=RedHatAI/Qwen3-Coder-Next-NVFP4'
- '--served-model-name=qwen3_coder'
- '--tensor-parallel-size=4'
- '--tool-call-parser=qwen3_coder'
- '--enable-auto-tool-choice'
- '--enforce-eager'
ports:
- containerPort: 8000
protocol: TCP
env:
- name: HF_HUB_OFFLINE
value: '0'
- name: HUGGING_FACE_HUB_TOKEN
value: <YOUR_HUGGING_FACE_TOKEN>YAMLThe base image names or runtime parameters may change. You can usually find the most up-to-date version in my repository, in the ai directory.
To expose the model outside of the OpenShift cluster, we apply the following to the YAMLs.
apiVersion: v1
kind: Service
metadata:
name: qwen3-coder
namespace: ai
spec:
selector:
app: qwen3-coder
ports:
- protocol: TCP
port: 8000
targetPort: 8000
---
kind: Route
apiVersion: route.openshift.io/v1
metadata:
name: qwen3-coder
namespace: ai
spec:
to:
kind: Service
name: qwen3-coder
weight: 100
port:
targetPort: 8000YAMLNow we need to be patient. It may take a few minutes to bring up the model. Eventually, however, you should see logs like the ones below. This means we’ve just launched Qwen3 Coder on Kubernetes!

Our model is accessible from outside the cluster at the address Route qwen3-coder. To check that URL on your cluster, run the following command:
oc get route qwen3-coder -n aiShellSessionIntegrate Custom Model with OpenCode on GitHub
Now that we’ve deployed the model externally, we can integrate it with OpenCode. First, define a new, custom OpenAI-compatible model in the GitHub repository. To do this, we’ll create a configuration file opencode.json as shown below. It must contain the provider name, the URL of the inference server, and the model name. The model name must match the name we set in its Deployment under the --served-model-name parameter. The provider name, on the other hand, can be anything, but we’ll use it later when calling the model from the GitHub application.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"vllm": {
"npm": "@ai-sdk/openai-compatible",
"name": "Qwen3 Coder Next",
"options": {
"baseURL": "http://qwen3-coder-ai.apps.piomin-aws-main.5fuc.p1.openshiftapps.com/v1"
},
"models": {
"qwen3_coder": {
"name": "Qwen3 Coder Next"
}
}
}
}
}JSONNow, we can go back to defining our GitHub workflow. It will be simpler than before for the Google Vertex AI-based model, because our endpoint does not require authorization. In the Run OpenCode section, set the provider name and the model name.
name: opencode
on:
issue_comment:
types: [created]
pull_request_review_comment:
types: [created]
jobs:
opencode:
if: |
contains(github.event.comment.body, ' /oc') ||
startsWith(github.event.comment.body, '/oc') ||
contains(github.event.comment.body, ' /opencode') ||
startsWith(github.event.comment.body, '/opencode')
runs-on: ubuntu-latest
permissions:
id-token: write
contents: read
pull-requests: read
issues: read
steps:
- name: Checkout repository
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Run opencode
uses: anomalyco/opencode/github@latest
with:
model: vllm/qwen3_coderYAMLNext, if you add a comment starting with /opencode to any issue in your GitHub repository, the AI agent will communicate with the custom model. In this case, you should see a communication trace in the qwen3-coder Pod logs.

OpenCode for Pull Requests Review
Introduction
In this section, we’ll analyze a slightly more advanced scenario involving code changes. In addition to implementing the changes based on the issue description, I’d like OpenCode to review the pull request that was created. This could be a pull request created by a developer or by a bot, such as Renovate or OpenCode Agent. For example, we can use a different model to review a pull request than to implement the task itself.
To do this, we’ll create additional GitHub Actions workflows. But first, I’d like to flag an OpenCode issue that will affect how we implement these workflows. You can read more about it in this OpenCode repository (as well as in a few others): https://github.com/anomalyco/opencode/issues/17224.
In short: if a pull request was created by a bot rather than a user, reviewing it using the OpenCode agent will simply result in a permission error. Perhaps OpenCode will resolve this issue at some point. For now, I’ll suggest a workaround. Our workflow will check the username of the user who created the pull request. It will then handle the pull request differently depending on whether a bot or a standard user created it.
Create and Run GitHub Review Workflows
The GitHub workflow visible below runs the OpenCode review for a standard user. For a bot user, it calls another workflow using the GitHub CLI.
name: Trigger OpenCode review
on:
pull_request:
types: [opened]
jobs:
trigger:
if: endsWith(github.event.pull_request.user.login, '[bot]')
runs-on: ubuntu-latest
permissions:
actions: write
steps:
- name: Trigger OpenCode Review [Bot]
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
gh workflow run opencode-review.yml \
--repo "$REPO" \
--ref "${{ github.event.repository.default_branch }}" \
-f pr_number="$PR_NUMBER"
human-pr-review:
if: "!endsWith(github.event.pull_request.user.login, '[bot]')"
runs-on: ubuntu-latest
permissions:
id-token: write
contents: read
pull-requests: write
issues: write
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- name: Create Google credentials
env:
GCP_SERVICE_ACCOUNT_JSON: ${{ secrets.GCP_SERVICE_ACCOUNT_JSON }}
run: |
set -euo pipefail
printf '%s' "$GCP_SERVICE_ACCOUNT_JSON" \
> "$RUNNER_TEMP/gcp-credentials.json"
- uses: anomalyco/opencode/github@latest
env:
GOOGLE_VERTEX_PROJECT: ${{ secrets.GOOGLE_VERTEX_PROJECT }}
GOOGLE_VERTEX_LOCATION: ${{ secrets.GOOGLE_VERTEX_LOCATION }}
GOOGLE_APPLICATION_CREDENTIALS: ${{ runner.temp }}/gcp-credentials.json
GITHUB_TOKEN: ${{ github.token }}
with:
model: google-vertex-anthropic/claude-opus-4-6@default
use_github_token: true
prompt: |
Review this pull request for:
- potential bugs
- correctness issues
- security problems
- code quality issues
- problematic edge cases
- unnecessary complexity
Focus on actionable findings that are relevant to the
changes in this pull request.
Post the review findings.opencode-review-trigger.ymlHere, in turn, you can see the workflow responsible for reviewing a pull request created by a bot. It was triggered by opencode-review-trigger.yml. As you can see, this workflow uses the pull request number. It checks out the pull request branch and then reviews the changes using the OpenCode agent, which communicates with the vllm/qwen3_coder model.
name: OpenCode review
on:
workflow_dispatch:
inputs:
pr_number:
description: "Pull request number"
required: true
type: string
jobs:
review:
runs-on: ubuntu-latest
permissions:
id-token: write
contents: read
pull-requests: write
issues: write
steps:
- name: Get PR information
id: pr
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ inputs.pr_number }}
REPO: ${{ github.repository }}
run: |
set -euo pipefail
echo "number=$PR_NUMBER" >> "$GITHUB_OUTPUT"
gh pr view "$PR_NUMBER" \
--repo "$REPO" \
--json url,title,headRefName,baseRefName \
> "$RUNNER_TEMP/pr.json"
cat "$RUNNER_TEMP/pr.json"
- name: Checkout PR
uses: actions/checkout@v7
with:
ref: refs/pull/${{ inputs.pr_number }}/head
persist-credentials: false
- name: Run OpenCode
uses: anomalyco/opencode/github@latest
env:
GITHUB_TOKEN: ${{ github.token }}
with:
model: vllm/qwen3_coder
use_github_token: true
prompt: |
Review pull request #${{ inputs.pr_number }}.
Repository: ${{ github.repository }}
Review the checked-out pull request for:
- potential bugs
- correctness issues
- security problems
- code quality issues
- problematic edge cases
- unnecessary complexity
Focus on actionable findings that are relevant to the
changes in this pull request.
Post the review findings to pull request #${{ inputs.pr_number }}.opencode-review.ymlIn general, I’m switching many of my repositories to a maintenance model using the OpenCode app on GitHub. That’s why you’ll easily find examples of issues and pull requests with automatic reviews performed by the OpenCode agent. You can see an example pull request below. It’s located in the https://github.com/piomin/sample-graphql-microservices repository. In such a pull request, CircleCI pipelines are triggered to perform the build and scan the code in SonarCloud. Additionally, a GitHub action is triggered that reviews the changes made by the agent itself 🙂

Here you can see the result of such a review.

Conclusion
OpenCode, GitHub Actions and Kubernetes show how AI can become a natural part of the development workflow — from writing code to reviewing pull requests. Running custom models on Kubernetes gives us even more flexibility: we keep control of our data, models, and infrastructure, while scaling AI workloads as needed. In practice, this turns the AI model into another internal service we can plug into our development tools and adapt to our needs.



Related Posts