Tech Blog
Auto-aggregated global tech articles · 1,047 posts
How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]
I’m the author of a new preprint on repeated-query auditing of LLM brand recommendations, and the founder of Rankfor.AI. The practical question: how many times should we repeat a prompt before comparing results? The paper applies generalizability theory: estimate variance components from a pilot, th
Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
Article URL: http://www.techdirt.com/2026/09/03/hackers-had-a-live-feed-of-every-id-this-verification-company-scanned-for-over-a-year/ Comments URL: https://news.ycombinator.com/item?id=49561320 Points: 35 # Comments: 7
fastdb version just got updated to v1.0.8
FastDB is an (persistent) in-memory key/value store in Go. I made some updates to comments I've got (A big thank you to ShotgunPayDay) Test coverage : 96.54% https://codeberg.org/marcelloh/fastdb A test does the following: - Populate 1,000 items with rdom.Intn(1000000). - Inside for range b.N, exec
I picked MongoDB on day one. By month eight I needed Postgres.
The decision took about nine minutes. It was the start of a new product — VirtualRx, a business-management app for small retailers in Lagos. Point of sale, inventory, invoices, the lot. I needed a database, I had a schema roughly in my head, and MongoDB let me start writing code that afternoon inst
Agent task length doubles every seven months. The reliable version is 18 months behind.
METR has been tracking one number since 2019, and it is not a benchmark score. It is a duration: the length of a task, measured in how long a human expert needs, that a frontier agent can finish on its own. The headline result from their NeurIPS 2025 paper is that this duration has been doubling ro
What Works and What Doesn't in CLAUDE.md
Writing "Write clean code" in CLAUDE.md changes nothing. During the process of building 10 personal apps in three months, I rewrote CLAUDE.md many times. Since it became clear what worked and what didn't, I will outline that distinction. What Doesn't Work Giving instructions with
Stop Building AI Agents. Start Building AI Systems.
There's a phrase I keep seeing everywhere in AI development: "We need an AI agent." Need to analyze documents? Build an agent. Need to automate a workflow? Build an agent. Need to interact with APIs? Build an agent. Need to write code? Build an agent. At some point, I started asking a diff
I Didn't Touch My Site for a Month. Traffic Was Stable — Then It Suddenly Dropped
I hadn’t really touched one of my AI projects, Vivify Face Swap, for about a month. No major updates, no new campaigns, no big SEO changes. Traffic stayed surprisingly stable the whole time. Then around August 30, it suddenly dropped. That was the interesting part for me. Nothing obvious chan
I built a jungle survival browser game where you co-op with an AI agent
This is my first post here. I figured I should start somewhere, and a hackathon project is as good a reason as any. The background I have been writing business software for about three years. Dashboards, APIs, data pipelines. It pays the bills but it stopped being fun a while ago. When I
Stop Paying for Vercel: How to Self-Host Your Deployments on a VPS and Keep the Same Workflow
Push-to-deploy, automatic HTTPS, and preview environments on a $24 droplet with Dokploy, plus the four mistakes to skip on the way. For two years, deploying was the easiest part of my week. I pushed to main, closed my laptop, and the site updated itself. I never thought about servers, which was the
Your offline GitHub
Hi all! I’ve been coding professionally for the last 15 years - and more and more with AI next to me. Naturally pushed by the market to produce more and more output, I now find myself hopping daily between 5 to 7 simultaneous AI sessions. I realized that the bottleneck is becoming the code review