PROMPTWIRE

There is a folder on my machine I have never opened in front of an AI.

Contracts with real numbers in them. A medical thing I wanted a second opinion on and did not want sitting on someone's server. Payroll. A dispute with a supplier where naming them in a chat window felt like a bad idea.

You have a version of that folder. Everyone does. And the standard advice is to be careful what you paste, which is not advice, it is just a warning with no solution attached.

Last week I stopped working around it. I installed a model on my own laptop. No account, no subscription, no internet connection required after the download. I asked it the medical question with the wifi switched off.

Took about five minutes to set up. That is the issue this week.

⚡ This Week In One Minute

  • Workflow: Run a real AI on your own machine. No account, no subscription, nothing leaving your device. Five minutes, and the setup is below.

  • Big Story: The chip in your pocket quietly became powerful enough to run serious models locally. What that changes about who sees your data.

  • Tool of the Week: Wispr Flow. Talk instead of type, and get finished text rather than a messy transcript.

  • New in the library: The Connector Playbook. 30 connectors for Claude and ChatGPT, with a copy-paste prompt for each.

🔁 This Week's Workflow

Put a Real AI on Your Own Machine

The 20-second version: Download one app, pick a model, click load. You are now chatting with an AI that runs entirely on your hardware. No account, no subscription, no data leaving the device, and it works with the internet switched off. The trade is that it is not as clever as Claude or ChatGPT on hard reasoning. For sensitive work, that trade is worth making.

You need: a laptop with 8GB of RAM. 16GB makes it noticeably better. Payoff: somewhere to put the work you have been carefully keeping out of chat windows.

The old way: you self-censor. You anonymise the contract before pasting it, or you strip the numbers out of the spreadsheet, or you just do that bit by hand because it is not worth the risk. Most people do not even notice they are doing it.

The replacement: the model runs on your machine. The weights and your prompts never leave the device, so there is nothing on a server to leak, subpoena, or train on.

Set It Up (5 Minutes)

Use LM Studio. There are more powerful options, and I will cover those in the Pro section, but this is the one with a proper graphical interface and no terminal, so anyone can run it (AI Thinker Lab).

  1. Download it from lmstudio.ai and install

  2. Use the search to find a model. Start with something in the 7B to 8B range

  3. Click download, wait for it, then click load

  4. Chat

That is the entire process. The model file is a few gigabytes, so the download is the slow part. Everything after that is instant and offline.

One thing that decides whether this feels good or terrible: when you pick a model, take the version tagged Q4_K_M or Q5_K_M. Those are compressed builds. The full-precision version eats your RAM for a barely noticeable quality gain. Most people who try local AI and conclude it is slow have skipped this.

The First Thing to Try

Do not test it with trivia. It will lose that comparison against ChatGPT and you will conclude it is useless.

Test it with the thing you have been avoiding.

I want your read on this and I want you to be direct rather than 
reassuring.

[paste the actual thing: the contract clause, the medical letter, the 
figures, the message you are unsure how to answer]

Tell me what stands out, what I should be worried about, and what 
question I should be asking that I have not asked.

What you get: a competent second opinion on something you would not have shown a cloud tool. Not a genius one. A competent one, on the actual document rather than a version you sanitised first.

Switch your wifi off before you send it. Not because you need to, but because watching it answer with the connection dead is the moment the whole thing clicks.

The Honest Limits, Up Front

I would rather you hear this from me than discover it and give up.

A model running on your laptop is small. Cloud services run models with hundreds of billions of parameters on data centre hardware. Yours is running something in the single-digit billions. For hard reasoning, long chains of logic, or anything needing current information from the web, the cloud is still clearly better.

So this does not replace Claude or ChatGPT. It sits beside them, and it takes the category of work you were never comfortable sending anyway.

That is a smaller claim than most coverage makes, and it is the true one.

Which Model, Which Job, and the Bit Nobody Explains

The setup is easy. The part that decides whether this is genuinely useful is the part nobody covers: matching the model to your hardware and your job, and then the feature that turns this from a chat toy into something that reads your actual files.

That is the Pro section this week.

🔒 Inside the Pro section:

  • The model-to-hardware table. What actually runs well on 8GB, 16GB, and 32GB, and the download that will crawl on your machine

  • Private document Q&A. Point it at a folder of your own contracts, records, or notes and interrogate them, entirely offline

  • The 5 jobs worth doing locally, each with the full prompt written out

  • The 4 prompting rules for small models, which are the reverse of what you have learned

  • The phone version. Yes, this runs on your phone, and yours may already have a model built in

  • The four mistakes that make people quit in the first hour

🔒 The Full Setup

🔐 This Week For Pro Members

The Private AI Pack. The setup above gets you running. This is how to make it genuinely useful:

  • The model-to-hardware table. 8GB, 16GB, and 32GB machines, what each actually handles, and the specific models worth downloading for each. Plus the download size trap that makes a good laptop feel broken.

  • Private document Q&A, step by step. The app that lets a local model read a folder of your own files and answer from them, offline. This is the feature that makes it worth the disk space, and it is the one lawyers and researchers use.

  • The 5 jobs worth doing locally, with the full prompt for each: the contract read, the financial sanity check, the difficult message, the medical or legal second opinion, and the one involving other people's names.

  • How to prompt a small model, which is the opposite of how you prompt Claude, and the reason most people conclude local AI is useless.

  • The phone setup. Which app, which model, and how to check whether your phone already has one built in.

  • The four quitting mistakes, including the quantization one and the reason your GPU might be sitting idle while you blame the model.

Also new in your library this week: The Connector Playbook. 30 connectors for Claude and ChatGPT with a copy-paste prompt for each, including the ones almost nobody has found.

🔧 Tool of the Week

You type slower than you talk. Everyone does, by roughly a factor of three, and yet almost all of us still write emails at typing speed.

Wispr Flow turns speech into clean, finished text. Not a raw transcript with your ums and false starts in it, but writing you can send. It works for emails, documents, notes, messages, and, usefully, for writing AI prompts.

The prompt use is the one I did not expect to care about. Long, detailed prompts are what get good output, and long detailed prompts are tedious to type, so most people write short lazy ones instead. Talking a full brief takes twenty seconds.

🔥 This Week in AI

📰 Short Updates

🎙️ An AI notetaker exposed 181,874 meetings, including calls that were live at the time. A researcher found tl;dv's database had no separation between customers, exposing records from 84,312 users across 35,003 domains, with around 1,000 meetings showing as actively recording at any moment (Dark Reading). Default your meeting tools to private and you are covered even when the vendor is not.

🧩 A capable open model landed that runs on ordinary hardware. Qwen released Qwen3.8-27B on August 14, the most recent tracked frontier release at time of writing (AI Release Tracker). Open-weight releases are what make this week's workflow possible. Every one of them means better local options on the machine you already own.

💸 AI got dramatically cheaper again. OpenAI cut GPT-5.6 Luna pricing by 80% to $0.20 per million input tokens on July 30, aimed squarely at high-volume automation (AIapps). Anything you costed out earlier this year as too expensive to run at volume is worth pricing again.

🎬 A voice memo now becomes a finished video. VizNow launched Idea to Video on August 14, turning an idea, a script, or a voice recording into a complete video in one workflow (Agentic.ai). If you have been meaning to make video and never got past the blank timeline, talking for three minutes is now a viable starting point.

📖 Big Story of the Week

The Chip in Your Pocket Got Serious

Quick version: The newest phone processors ship neural engines able to run genuinely large models entirely on the device. That means a capable assistant that knows a great deal about your life, with a hard guarantee that none of it reaches a company server. The bottleneck for AI stopped being raw computing power and became electricity and memory. What that changes is not speed. It is who sees your data.

For three years the deal has been fixed. You get frontier intelligence, and in exchange your questions go to a data centre owned by someone else. Everybody accepted it because the alternative was nothing.

That deal is slowly coming apart. The current generation of mobile chips, including Apple's A20 Bionic and the latest Snapdragon, carry neural processing units capable of running 70-billion parameter models entirely on-device, which puts a genuinely capable assistant on your phone with no cloud connection at all.

The context that made this urgent rather than merely interesting arrived the same month. A single AI notetaking company exposed 181,874 meetings, including ones recording at that moment (Dark Reading). Nobody using that tool did anything wrong. They trusted a vendor, which is the only option the cloud model offers you.

Local changes the question. Instead of "do I trust this company with this," it becomes "does this need to leave my device at all." For a large share of what most people do with AI, the honest answer is no.

There is a second force pushing the same direction, and it is unglamorous. Running data centres for cloud AI takes an enormous amount of power, and the industry's real constraint now is electricity and memory rather than compute (TechDG). Every query answered on a device is a query nobody has to power in a warehouse.

So privacy and economics are pointing the same way for once, which is usually when something actually happens.

The part worth thinking about is what you would do differently if the trust question disappeared.

🔒 The Full Breakdown

🔒 The Pro breakdown: the work you are already self-censoring without noticing, the split worth making between local and cloud, and what to do about the vendors already holding your data.

📦 New Resources Added

New This Week 🚀

  • The Private AI Pack: the model-to-hardware table, private document Q&A on your own files, the 5 jobs worth doing locally, the phone setup, and the four mistakes that make people quit.

  • What Changes When Nothing Leaves: the self-censoring audit, the local versus cloud split, and the vendor cleanup worth an hour this month.

  • The Twitter Growth Skill

  • The Numbers Skill

  • The Marketing Psychology Skill

  • The Connector Playbook: 30 connectors for Claude and ChatGPT with a copy-paste prompt for each, the setup path for all four connector types, the power combos, and the ones almost nobody has found yet.

Each resource lives permanently in your Pro account. Use them whenever you need them.

Until Next Week

The interesting part was not the technology. It was noticing how much I had been working around a problem I had stopped seeing.

Install it while the kettle boils, switch your wifi off, and ask it the thing you would not ask anything else. Five minutes, no account, no subscription, and nothing to leak.

🔐 Why People Subscribe

👇 What’s behind the paywall:

  • The Private AI Pack: the model-to-hardware table, private document Q&A on your own files, the 5 jobs worth doing locally, the phone setup, and the four mistakes that make people quit.

  • What Changes When Nothing Leaves: the self-censoring audit, the local versus cloud split, and the vendor cleanup worth an hour this month.

  • Breakdown plan to implement AI into your workflow

  • Workflow Library with all past workflows, plus a new one every week

  • Resource bank built up of past resources

  • All past issue archives and walkthroughs

Pro members deploy AI in their work an average of 5-8x more often than free readers (based on reply data from past issues). The difference is having the exact setup, not the concept.

Till next time,

PROMPTWIRE