683 points by ilreb 1 day ago | 285 comments | View on ycombinator
mmastrac 1 day ago |
corysama 1 day ago |
https://news.ycombinator.com/item?id=49736660
https://www.reddit.com/r/LocalLLaMA/comments/1wjieap/made_th...
Papers: https://arxiv.org/abs/2503.23303 https://arxiv.org/abs/2510.01237
Model: https://huggingface.co/DeepMostInnovations/sales-conversion-...
Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sal...
wuhhh 1 day ago |
"Jev is TypeSafe's closed service for runtime-defined semantic decisions. This project reproduces that interface pattern with open models; it does not reproduce Jev's undisclosed model or training"
As someone else pointed out it isn't actually Jev... can someone enlighten me
prodigycorp 1 day ago |
kul_ 1 day ago |
lucfranken 1 day ago |
Also with this example the speed of new launches based on a launch is just incredible.
jakozaur 1 day ago |
Though Jev is original, it looks highly replicable.
kouteiheika 1 day ago |
brap 1 day ago |
We’ve always had output schemas for LLMs, and we’ve had small language classifiers for decades, so what’s new? Is it just some sweet spot in between in terms of quality vs speed?
wg0 1 day ago |
ritzaco 1 day ago |
druskacik 1 day ago |
If it was possible to re-create it as an open-weight, it would be exciting!
mohsen1 1 day ago |
vitonsky about 6 hours ago |
undefined about 6 hours ago |
perfectbeeing about 3 hours ago |
ludicrousskill 1 day ago |
2 answers: Yes No
- Qwen3 direct Read Yes: 0.985 No: 0.015 - Qwen3 generation Yes: 0.5 No: 0.5
- MiniCPM5 direct read Yes: 0.122 No: 0.878 - MiniCPM5 generation Yes: 0.5 No: 0.5
- Qwen3.5 direct Read Yes: 0.529 No: 0.471 - Qwen3.5 generation Yes: 0.95 No: 0.05
I feel we're just getting coinflip answer faster.
algoth1 1 day ago |
dev_l1x_be about 9 hours ago |
dankobgd 1 day ago |
jFriedensreich about 11 hours ago |
mukundesh 1 day ago |
hmokiguess 1 day ago |
tomaytotomato 1 day ago |
Are there any huggingface mirrors out there?
aatd86 about 24 hours ago |
tmach32 1 day ago |
I think one difference between OpenJev and Jev would be, then, is what it's trained on.
Jev is, on the surface, cheap enough for me not to seek self-hosted alternatives. On the other hand, I wish the free/open weight alternatives to Pangram were better.
paulluuk 1 day ago |
Probabilistic: 1.968 s - 76% chance it lands on a 1.
Generation: 3.083 s - Equal split.
neilellis 1 day ago |
tantalor 1 day ago |
hbarka about 23 hours ago |
khalidx 1 day ago |
yuppiepuppie about 9 hours ago |
daxaxelrod 1 day ago |
stpedgwdgfhgdd 1 day ago |
or it is just incredible slow - and I picked the smallest model…
Refreshing, model still in cache, but did not help.
phoghed 1 day ago |
As opposed to a fake choice?
brunooliv 1 day ago |
manerMon1 1 day ago |
singularity2001 1 day ago |
rogerdickey 1 day ago |
"after seeing the ghost he was sh*tting bricks"
is this person: pooping? 95% scared? 5%
:)
cmrdporcupine 1 day ago |
The thing is that the openjev stuff is a ... bit ... of a hack (a good one though):
It does this:
1. Send a throwaway request containing the shared state.
2. Hope SGLang keeps that text in its prefix cache.
3. Send a separate request for every question.
4. Each request repeats the shared beginning (but SGLang hopefully reuses the cached work in.)
5. Compute the complete vocabulary ; hundreds of thousands of possible tokens.
6. Keep only the few special answer tokens.
7. Convert those scores into probabilities.
Obviously this can all be done way more elegantly if you just own the inference engine -- fork / modify SGLang or vllm or llama.cpp, or do what I did in my bespoke inference engine (https://github.com/rdaum/eider/ commit https://github.com/rdaum/eider/commit/b2f981b7ebe0e338f60188...)
that ends up being, instead:
1. Convert the state into one shared prompt.
2. Run that shared prompt through the model once.
3. Fork the model’s internal state once per question.
4. Add a different question to each fork.
5. Ask each fork for its next-token scores.
6. Calculate only 64 possible label scores—not the whole vocabulary.
7. Convert the relevant scores into probabilities and return structured JSON.
I expect we'll see patches for llama.cpp and the others over the next few days/weeks and I also expect most model hosting providers will just end up providing this same service. I don't think Jev themselves have much of a moat. Though maybe it's more about their specific model and the training it gets.
FooBarWidget 1 day ago |
tirtha 1 day ago |
tecleandor 1 day ago |
It's trying to "emulate" Jev behavior using a regular small LLM model (Qwen3 0.6B or MiniCPM5 2B). And with the smallest model it takes like between half to two seconds to run in my M2 Max, so it's not super fast.
I mean, it's faster than asking to a regular LLM, but I think that's not proper to have Jev on the name (also legally...)
Edit: no shade, and I'll give it a try for some ideas. I'd also like to have an open weights Jev but I think the naming is misguiding. I also have to try Jev that, BTW, got access pretty quickly, less than a day I think...
speedgoose 1 day ago |
Between "brocoli and poop soup" or "cake", it recommends me to eat the soup.
jasurme 1 day ago |
zemlyansky 1 day ago |
AIorNot about 23 hours ago |
bnbn88 about 18 hours ago |
exe34 1 day ago |
spwa4 1 day ago |
The last step of an LLM is to take a softmax of the predictions and then generating a token from that. But there was tooling that would just generate all allowed next tokens from a grammar (e.g. restrict to valid JSON), zeroing all the ones not allowed and then picking the best among the allowed tokens.
This seems to taking an approach from the pre-transformer days. Seq-to-seq is hard and we don't always need it. So let's do seq-to-1 because it's often way easier to get it training properly and so you can often get it optimized way better. And, more generally, make sure to pick the best option out of the possibilities: 1-to-1, 1-to-seq, seq-to-1 and seq-to-seq. Where seq-to-seq requires far more resources than any other option and so it's a case of "please don't".
Also note that "1" only means the input is fixed. It does not mean 1 number or ... it just means fixed. The best image description models remained 1-to-seq models 4 years or so after transformers were introduced. Even ASR models remained 1-to-seq + CTC to stitch overlapping parts together to a final prediction ... I'm not sure if they lasted all the way to whisper release.
Even today training transformers remains expensive. So this should at least be a way to be a lot cheaper than any LLM can hope to be.
And I really like the doom demo. Obviously a pretty stupid model which is really cheap to run can still get a robot walking, if you run it quickly enough. That's how we get insects and mice and ...
And one might even add that biologically, humans aren't smart, or at least, most of the human nervous system isn't smart, compared to the whole, and does work independently if needed (and possible). The human mind is a LOOOOOOOONG chain of fast-but-stupid-and-totally-blind -> slightly-slower-but-smarter-and-not-entirely-blind -> slower-smarter-and-actually-senses-things -> all-information-you-could-want-but-at-most-1-signal-per-minute. We have "neural circuits" (using Bishop's definition) that can run at >2khz (2000+ tok/s, say, but you probably can't teach anything more than averaging) and on the other end up to our frontal lobe that takes one decision per week if it feels like working hard, and seems to decide on it's prediction of the future weeks to months out. Months or years if you're 40 or older.
camillomiller 1 day ago |
"Customer wants to lear how to better talk in a company situation, and bring across their argument effectively"
Than had it choose what training would be fitting for this user: - Communication and Feedback - Leadership for Begninners - Soft Skills and Emotional Awareness
It picked always the third with an 80% confidence, while the answer should have been 1.
ares623 1 day ago |
colesantiago 1 day ago |
Learned also that Jev was trained on 100%(!) synthetic data.
What a great time to be alive.
shying 1 day ago |
baobabKoodaa 1 day ago |
airza 1 day ago |
hbcdbff 1 day ago |
jamesforestwest about 20 hours ago |
estetlinus about 13 hours ago |
theoleecj 1 day ago |
bikeshedder2 about 22 hours ago |
On my DGX Spark I get very similar latency numbers, and it matches my evals + or - a few points on each test (DG wins some, Jev wins some, both show low confidence when wrong).
I ran the same evals against a Qwen36 and it clearly lost to both of them, so you are leaving both knowledge and instinctual reasoning on the table with any smaller models, FWIW.
https://github.com/vllm-project/vllm/pull/57250