Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

From the article:

  I've been seeing some very promising results from DeepSeek R1 for code as well. Here's a recent transcript where I used it to rewrite the llm_groq.py plugin to imitate the cached model JSON pattern used by llm_mistral.py, resulting in this PR.
But the transcript mentioned was not with Deepseek R1 (not the original, and not even the 1.58 quantized version), but with a Llama model finetuned on R1 output: deepseek-r1-distill-llama-70b

So perhaps it's doubly impressive?



Yeah, I was using the lightning fast Groq-hosted 70B distilled version.


Did you happen to try the same thing on Deepseek R1 on https://chat.deepseek.com/ ?


No. I tried it just now with the same prompt and got a similar looking response (with some different design decisions but I'd expect that for even the exact same model). https://gist.github.com/simonw/115620647028336e3a1edfe8a48e1...




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: