Language models for text classification: From bag-of-words to Jev

(magazine.sebastianraschka.com)

202 points | by Anon84 1 day ago

11 comments

  • malshe 11 hours ago
    Sebastian is the author of two excellent books related to LLMs

    Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/

    Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/

  • Topfi 16 hours ago
    Solid assessment, very much what I assumed after announcement.

    The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.

    Hype really is the worst aspect of this industry.

  • nzoschke 1 day ago
    Great article. This matches my feelings:

    > Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.

    We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email

  • aesthesia 8 hours ago
    I like the way this builds up from simple models to transformers. There's still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum/average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that's analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
  • kevinwang 11 hours ago
    Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.
  • jacek-123 17 hours ago
    Thats very nice, I always explain LLMs to ppl starting from old-school language models and then just replacing the predictor from bag-of-words to transformers to etc.) I think its a cool framing
    • alansaber 16 hours ago
      Agreed as well, this is my preferred framing, it only makes sense to get into semantic vectors once you realise the very real limits of discrete vectors and bag of words.
  • tomrod 23 hours ago
    Dr. Raschka is wonderful. Great article.
  • frjj 14 hours ago
    [dead]
  • alescalaios 17 hours ago
    [flagged]
  • Moon_Y 20 hours ago
    Really enjoyed the historical framing here. Starting from bag-of-words and working up to Jev makes the “just a classifier” argument much more nuanced. The tradeoff between generality, speed, and cost is probably the most interesting part to me — especially where Jev sits between task-specific classifiers and full LLMs.
  • aidiscoverywire 21 hours ago
    The calibration section is the most important part of this article for anyone actually shipping classifiers. nzoschke's email use case is exactly where this bites: for production routing, you almost always want to threshold on confidence ('auto-handle above 0.9, route to a human below'), and that only works if the probabilities mean what they say. A 96% model with overconfident outputs is operationally worse than a 94% model with honest ones. The Guo et al. observation Raschka cites — networks can overfit NLL without overfitting 0/1 loss — is why this happens with plain fine-tunes, and it's also the strongest part of Jev's design: training the confidence directly (RLCD/RLCR-style) rather than bolting calibration on after the fact. One caveat I'd add to the IMDb numbers: as the article notes, we don't know whether IMDb was in Jev's synthetic training mix. Until someone runs these benchmarks on a private, never-published dataset, take the 96% as a ceiling, not a measurement.