Ask HN: How are you getting inference for personal projects?

It seems like the major providers are starting to cut back the usage included with their monthly subscriptions. How are people handling inference for side projects these days?

Are you still using Codex/Claude subscriptions, paying for hosted open weight models, running local models, etc

4 points | by variety8675 3 days ago

9 comments

  • AlisaYoki 3 hours ago
    I use Cursor, and in fact, their internal models are more than enough for me. Especially considering that Grok bot has recently been released; it has its own limits and also helps save Cursor tokens (by properly drafting technical specifications and making edits to tasks). I pay $60 per month for the subscription, sometimes $200 if there are a lot of technical tasks
  • iliedanila 3 days ago
    In order to have an efficient use of tokens, I created proper documentation from the start of the project: - What component is where, what it does and how it does it - Kept a history of all the architectural decisions - How a new component (of a known type) should be added to the system - Instructed the agent to keep this information valid and up to date as part of the PR prerequisites.
  • spottedmarley 3 days ago
    I've been on the Claude Max plan for quite a while now but I also run my own local models and every week there are new local models that are slowly but very surely closing the gap to frontier-level quality and so, pretty soon, I won't use APIs at all. Can't wait!
    • xpnsec 1 day ago
      Out of interest, what does your local setup look like? Is this a spare GPU you’re using or have you got something more designed for local LLMs?
  • intrr 3 days ago
    Claude subscription (the cheapest $20/month one) -- I very rarely hit the limits, particularly the weekly ones. But then again, the largest codebase I ever work on is 50k lines.

    I'd prefer local models massively for both work and "pleasure" (I love LLMs as random cognitive sparring partners), but then again, my bank account massively disagrees :/

  • sshussain270 1 day ago
    I am using a tool I built which gives me a pre-flight estimate about token usage for different models, for a given task. It looks like this https://caspian.md/assets/feature-4-ai-models.mp4
  • sds357 3 days ago
    I have a Proxmox machine with dual P40 gpus that pass thru to a vm with ollama running for local inference. Been working great for over 2 years.
  • wangxiaoxiang55 1 day ago
    [dead]
  • ashasik454 3 days ago
    [flagged]