Paolo Perrone — Shipping Production AI: Agents, Inference, GPU. Read by 1M+ AI engineers.
Shipping Production AI: Agents, Inference, GPU. Read by 1M+ AI engineers.
Paolo Perrone ranks #64 of 18,566 LinkedIn creators in Computer Software, and is a standout voice in United States. They have 134.5K followers and published 54 posts in the last 60 days at a 0.2% average engagement rate.
- 134.5K followers
- 54 posts / 60d
- 0.2% avg engagement
- 3.3K follower growth / 30d
The roast
Paolo claims he writes about how production AI actually works, yet he’s spent 50 posts in the last month explaining the industry to 131,000 people who are clearly only following him to see how much faster a career can be vaporized by an engagement rate of 0.15%.
About Paolo
Get your AI product in front of 1M+ engineers buying AI infrastructure, agent tooling, and inference stacks. NVIDIA, Google, LangChain, CodeRabbit already have. 📮 https://tally.so/r/VL4qBj (or shoot me a DM) I write about how production AI actually works. Agents, inference, GPUs, retrieval, evals.🎯 What I cover:→ Agents: orchestration, memory, planning loops, where they break→ Inference: latency, throughput, batching, serving (vLLM, TGI, TensorRT)→ GPUs: CUDA, kernels, model parallelism, what's fast on H100/B200→ Retrieval: vector DBs, hybrid search, chunking strategies→ Evals: measuring model performance when there's no ground truth📣 Distribution:→ LinkedIn: 130K+ AI/ML engineers→ Medium: 922K+ via Data Science Collective (Founding Editor)→ The AI Engineer newsletter: 25K+ AI engineers getting dangerously good at AI→ The Tech Audience Accelerator: 14K+ AI/ML founders building serious tech audiences🧑🏻💻 Background: → 8+ years shipping production ML systems→ Credit scoring, churn prediction, fraud detection. → My ML consultancy got acquired in 2024. → Now I write code AND content full-time. 👻 For AI/ML founders: No one funds the founder no one's heard of. I build the LinkedIn presence that pulls in investors, customers, and senior hires. 3-4 founders at a time 📮 https://tally.so/r/wk69VJ (or shoot me a DM) ⚡ For AI/ML companies: 1M+ engineers read my work. 100+ AI/ML companies sponsor to reach them: NVIDIA, Google, CodeRabbit, MongoDB. Solo operation. Running on the agent infrastructure I built. 📮 https://tally.so/r/VL4qBj (or shoot me a DM) P.S. Engineer first. Writer secondBest of both worlds ✌️
Highlights
- Top 1% in Computer Software — Ranked #8 of 4742 creators
- Top 1% Audience — 134,513 followers
- Top 1% in United States — Ranked #23 of 5851 creators
- Top 1% Creator — 54 posts in 30 days
Recent posts
I stopped using Obsidian. My notes live in the terminal now. shiki is a three-pane, Rust-written note-taking app that never leaves the terminal. 𝗪𝗵𝗮𝘁 𝗺𝗮𝗸𝗲𝘀 𝗶𝘁 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁: → Notebooks are actual git repositories, not a proprietary sync format → Per-note version history you can browse and revert from the TUI → 100% keyboard, fully remappable (Vim-style navigation) → Point it at your existing Obsidian vault → no migration, no symlinks → 15 built-in themes (Catppuccin, Tokyo Night, Gruvbox, Nord, Dracula) No database. No proprietary format. No vendor lock-in. If your term
40 reactions · 9 comments · 0 reposts
OpenClaw eliminated cloud AI dependencies. Three commands. One local machine. Zero API calls. git clone → ./install.sh → ./start_all.sh Your agent runs on your hardware. Reads your files. Never sends a byte to someone else's server. What you get: 1️⃣ No token tax Cloud agents cost $500-2K/month in API fees. OpenClaw costs hardware upfront, then $0/month. The math flips after 3-6 months. 2️⃣ 200B parameter models on DGX Spark 128GB unified memory. Full-size reasoning models. No quantization compromises. Requires NVIDIA DGX Spark (~$4K) or similar hardware. 3️⃣ 128K context. No rate limi
6 reactions · 1 comments · 0 reposts
Quantization always meant one thing: smaller model, dumber output. Google's TurboQuant just broke that rule. 1️⃣ Polar coordinate compression Think GPS vs street address. Same location, different format. Standard models store weights as XYZ coordinates. TurboQuant converts to radius + angle → more compressible, less memory. 2️⃣ Johnson-Lindenstrauss transform Compresses data from thousands of dimensions to single bits: +1 or -1. The trick: distances between points survive the compression. You lose raw precision. You keep what matters: relationships. → Polar compression shrinks storage. →
8 reactions · 4 comments · 0 reposts
Last updated 2026-08-01