Jump to content

AI LLM for Coding: Difference between revisions

From Torben's Wickie
mNo edit summary
 
(62 intermediate revisions by the same user not shown)
Line 1: Line 1:
==Mistral Devstral==
 
==My Current Setup==
* OpenCode with free default model
* Code formating via Ruff (python) and biome (node)
* 6 lines of Cavenman in global user home AGENTS.md
* 1 section of Ponytail in global user home AGENTS.md
* Context7 as MCP server for OpenCode
* RTK for token reduction
 
Recently, I moved all of it into a docker container to prevent LLMs from installing packages on my computer and having full access to all my files.
see https://github.com/entorb/devbox
 
=== Code checks===
See
* [[Code Checks]]
* [https://github.com/entorb/korrekturleser/tree/main/scripts Korrekturleser Repo]
 
==Tips==
* E2E tests (e.g. Cypress/Playwrigth) are very important
* Use tools for code formatting, linting, type-hints, and dead code
* Configure these tools to produce minimum output to reduce token consumption. e.g. for
** Prettier: --log-level silent
** Cypress: cypress run --e2e --quiet
* For Refactoring/renaming: better do it manually, via VSCode, that is a lot faster
 
===Caveman speech===
less words = less tokens
 
Accoring to [https://medium.com/@KubaGuzik/i-benchmarked-the-viral-caveman-prompt-to-save-llm-tokens-then-my-6-line-version-beat-it-d8e565f95e15]
Caveman speech can be can be reduced to 6 hard coded lines in (global) AGENTS.md:
Respond like smart caveman. Cut all filler, keep technical substance.
- Drop articles (a, an, the), filler (just, really, basically, actually).
- Drop pleasantries (sure, certainly, happy to).
- No hedging. Fragments fine. Short synonyms.
- Technical terms stay exact. Code blocks unchanged.
- Pattern: [thing] [action] [reason]. [next step].
 
Original Repo: [https://github.com/JuliusBrussee/caveman]
npx skills add JuliusBrussee/caveman -a github-copilot
It has an option to install to local home dir instead of to project repo.
 
===Ponytail, lazy senior dev mode===
add contents of [https://github.com/DietrichGebert/ponytail/blob/main/.agents/rules/ponytail.md] to (global) AGENTS.md
 
===llmfit: Right-sizes LLM models to your system's RAM, CPU, and GPU===
https://www.llmfit.org offers a tool that suggests local LLMs that run on your hardware.
 
===Context7===
https://context7.com offers an MCP server of latest spec to many languages, enabling the LLM to support latest versions and best practice.
npx ctx7 setup
In prompt write:
use context7. do xxx
 
===RTK - Rust Token Killer===
https://www.rtk-ai.app reduces shell command output
brew install rtk
brew upgrade rtk
rtk init --global
rtk init --global --opencode
rtk init --global --gemini
rtk init --global --agent vibe
 
===Codeburn (not tested by me)===
https://github.com/getagentseal/codeburn
npm install -g codeburn
# or
brew install codeburn
 
==Coding Agents==
 
===OpenCode===
https://opencode.ai
brew install anomalyco/tap/opencode
# add context7 MCP
npx ctx7 setup --opencode
 
# initialize Opencode and create/update/improve/compress AGENTS.md file
  /init
 
# update
brew update
brew upgrade anomalyco/tap/opencode
# update models
opencode models --refresh
 
Config ~/.config/opencode/opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "share": "disabled",
  "formatter": true
}
 
Addons
# add Caveman and Ponytail instructions to global
# ~/.config/opencode/AGENTS.md
# see below
# add Ruff Token Killer
rtk init --global --opencode
 
====Model ranking====
Most prominent models: https://opencode.ai/data/
 
====SQLite DB====
~/.local/share/opencode/opencode.db
SELECT title,
DATETIME(time_created/1000, 'unixepoch') AS date, -- UTC
round(cost,2) AS cost,
tokens_input, tokens_output, tokens_reasoning, tokens_cache_read, tokens_cache_write,
json_extract(model, '$.id') AS 'model-id',
json_extract(model, '$.providerID') AS 'model-providerID'
FROM SESSION s
WHERE 1=1
--AND project_id = 'fe41f10cbae7d51e25088b6ed35c168da9d3a4f6' -- flashcards
AND directory LIKE '%/flashcards'
AND model LIKE '%deepseek%'
ORDER BY time_created DESC
;
 
===Mistral Devstral Vibe CLI===
# install
  uv tool install mistral-vibe
  uv tool install mistral-vibe
# setup --> choose automatic
vibe --setup
  # run via
  # run via
  vibe
  vibe


Obtain API key from [https://console.mistral.ai/codestral/cli]
to set Cavemen style global, create ~/.vibe/AGENTS.md (or link to existing file)
 
connect to context7 MCP server
~/.vibe/config.toml (copy from edit mode)
[[mcp_servers]]
name = "context7"
transport = "http"
url = "https://mcp.context7.com/mcp"
[mcp_servers.headers]
CONTEXT7_API_KEY = "XXX"


==Google Gemini==
===Claude===
pros: free to use with google account
====Web====
create your API key at https://aistudio.google.com/app/api-keys
Website: [https://claude.ai]
* Pros: Daily reset free quota
* Cons: Need to manually upload single files
 
====Claude Console====
Claude Console via API and VISA Pre-Paid: [https://console.anthropic.com]
* Pros: Fully integrated in codebase
* Cons: Costs money (e.g., 5€ are easily spent)
npm install -g @anthropic-ai/claude-code
cd myProject
claude
Uninstall
npm uninstall -g @anthropic-ai/claude-code
 
===GitHub Copilot===
global config file:
.copilot\copilot-instructions.md
 
===Google Gemini===
* Pros: Free to use with Google account
Create your API key at [https://aistudio.google.com/app/api-keys]


Install
Install
Line 15: Line 169:
  # or just run
  # or just run
  npx https://github.com/google-gemini/gemini-cli
  npx https://github.com/google-gemini/gemini-cli
usage
Uninstall
npm uninstall -g @google/gemini-cli
 
Usage
  cd /tmp
  cd /tmp
  export GEMINI_API_KEY="xxx"
  export GEMINI_API_KEY="xxx"
Line 22: Line 179:
  gemini
  gemini


to switch model from pro to flash:
To switch model from pro to flash:
  export GEMINI_MODEL="gemini-2.5-flash"
  export GEMINI_MODEL="gemini-2.5-flash"


secret alternative: .env file
Secret alternative: .env file
  GEMINI_API_KEY=xxx
  GEMINI_API_KEY=xxx


==Google Antigravity==
===Google Antigravity===
https://antigravity.google/download
https://antigravity.google/download
==Claude==
===Web===
https://claude.ai <br>
pro: Daily reset free quota<br>
cons: need to manually upload single files
===Console===
Claude Console via API and VISA Pre-Paid: https://console.anthropic.com <br>
pro: fully integrated in code base<br>
cons: costs money, 5€ are easily spend
npm install -g @anthropic-ai/claude-code
cd myProject
claude
==Tips==
* E2E tests (e.g. Cypress) are very important
* Use tools for code formatting, linting, type-hints
* Configure these tools to produce minimum output to reduce token consumption. e.g. for
** Prettier:  --log-level silent
** Cypress: cypress run --e2e --quiet
* For Refactoring/renaming: better do it manually, via vscode, that is a lot faster
===Cavemen style===
https://github.com/JuliusBrussee/caveman
npx skills add JuliusBrussee/caveman -a github-copilot
It has an option to install to local home dir instead of to project repo.

Latest revision as of 09:10, 4 October 2026

My Current Setup

  • OpenCode with free default model
  • Code formating via Ruff (python) and biome (node)
  • 6 lines of Cavenman in global user home AGENTS.md
  • 1 section of Ponytail in global user home AGENTS.md
  • Context7 as MCP server for OpenCode
  • RTK for token reduction

Recently, I moved all of it into a docker container to prevent LLMs from installing packages on my computer and having full access to all my files. see https://github.com/entorb/devbox

Code checks

See

Tips

  • E2E tests (e.g. Cypress/Playwrigth) are very important
  • Use tools for code formatting, linting, type-hints, and dead code
  • Configure these tools to produce minimum output to reduce token consumption. e.g. for
    • Prettier: --log-level silent
    • Cypress: cypress run --e2e --quiet
  • For Refactoring/renaming: better do it manually, via VSCode, that is a lot faster

Caveman speech

less words = less tokens

Accoring to [1] Caveman speech can be can be reduced to 6 hard coded lines in (global) AGENTS.md:

Respond like smart caveman. Cut all filler, keep technical substance.
- Drop articles (a, an, the), filler (just, really, basically, actually).
- Drop pleasantries (sure, certainly, happy to).
- No hedging. Fragments fine. Short synonyms.
- Technical terms stay exact. Code blocks unchanged.
- Pattern: [thing] [action] [reason]. [next step].

Original Repo: [2]

npx skills add JuliusBrussee/caveman -a github-copilot

It has an option to install to local home dir instead of to project repo.

Ponytail, lazy senior dev mode

add contents of [3] to (global) AGENTS.md

llmfit: Right-sizes LLM models to your system's RAM, CPU, and GPU

https://www.llmfit.org offers a tool that suggests local LLMs that run on your hardware.

Context7

https://context7.com offers an MCP server of latest spec to many languages, enabling the LLM to support latest versions and best practice.

npx ctx7 setup

In prompt write:

use context7. do xxx

RTK - Rust Token Killer

https://www.rtk-ai.app reduces shell command output

brew install rtk
brew upgrade rtk
rtk init --global
rtk init --global --opencode
rtk init --global --gemini
rtk init --global --agent vibe

Codeburn (not tested by me)

https://github.com/getagentseal/codeburn
npm install -g codeburn
# or
brew install codeburn

Coding Agents

OpenCode

https://opencode.ai

brew install anomalyco/tap/opencode

# add context7 MCP
npx ctx7 setup --opencode
 
# initialize Opencode and create/update/improve/compress AGENTS.md file
 /init
 
# update
brew update
brew upgrade anomalyco/tap/opencode

# update models
opencode models --refresh

Config ~/.config/opencode/opencode.json

{
  "$schema": "https://opencode.ai/config.json",
  "share": "disabled",
  "formatter": true
}

Addons

# add Caveman and Ponytail instructions to global
# ~/.config/opencode/AGENTS.md
# see below

# add Ruff Token Killer
rtk init --global --opencode

Model ranking

Most prominent models: https://opencode.ai/data/

SQLite DB

~/.local/share/opencode/opencode.db

SELECT title, 
DATETIME(time_created/1000, 'unixepoch') AS date, -- UTC 
round(cost,2) AS cost, 
tokens_input, tokens_output, tokens_reasoning, tokens_cache_read, tokens_cache_write,
json_extract(model, '$.id') AS 'model-id', 
json_extract(model, '$.providerID') AS 'model-providerID'
FROM SESSION s
WHERE 1=1
--AND project_id = 'fe41f10cbae7d51e25088b6ed35c168da9d3a4f6' -- flashcards
AND directory LIKE '%/flashcards' 
AND model LIKE '%deepseek%'
ORDER BY time_created DESC

Mistral Devstral Vibe CLI

# install
uv tool install mistral-vibe
# setup --> choose automatic
vibe --setup
# run via
vibe

to set Cavemen style global, create ~/.vibe/AGENTS.md (or link to existing file)

connect to context7 MCP server ~/.vibe/config.toml (copy from edit mode)

mcp_servers
name = "context7"
transport = "http"
url = "https://mcp.context7.com/mcp"

[mcp_servers.headers]
CONTEXT7_API_KEY = "XXX"

Claude

Web

Website: [4]

  • Pros: Daily reset free quota
  • Cons: Need to manually upload single files

Claude Console

Claude Console via API and VISA Pre-Paid: [5]

  • Pros: Fully integrated in codebase
  • Cons: Costs money (e.g., 5€ are easily spent)
npm install -g @anthropic-ai/claude-code
cd myProject
claude

Uninstall

npm uninstall -g @anthropic-ai/claude-code

GitHub Copilot

global config file:

.copilot\copilot-instructions.md

Google Gemini

  • Pros: Free to use with Google account

Create your API key at [6]

Install

npm install -g @google/gemini-cli@latest
gemini
# or just run
npx https://github.com/google-gemini/gemini-cli

Uninstall

npm uninstall -g @google/gemini-cli

Usage

cd /tmp
export GEMINI_API_KEY="xxx"
git clone https://github.com/entorb/meeting-meter
cd myProject
gemini

To switch model from pro to flash:

export GEMINI_MODEL="gemini-2.5-flash"

Secret alternative: .env file

GEMINI_API_KEY=xxx

Google Antigravity

https://antigravity.google/download