Tools & Tech
Flex Inference: 50% Off LLM Calls on Gemini, OpenAI, and Bedrock
Every major AI provider now offers half-price inference if you can tolerate a few extra seconds of latency. One parameter change. Same API. Here's how it works and why.
Notes and ramblings, typically about LLMs.
Every major AI provider now offers half-price inference if you can tolerate a few extra seconds of latency. One parameter change. Same API. Here's how it works and why.
Using Gemini's bounding box detection to get precise measurements when converting a screenshot to code. Plus how prompt caching and flex inference make the multi-pass approach surprisingly cheap.
Universal Commerce Protocol lets AI agents buy things. Here's how developers can monetize it and what store owners need to know.
How to configure SSH so Claude Code can run commands on remote servers
Your *.pages.dev URL is ranking in Google instead of your custom domain. Here's how to fix it with a proper 301 redirect.