橘橘橘子汁 & 🍊 头像

消息来源频道

橘橘橘子汁 & 🍊

@microblock_pub

频道13,581 位成员公开可见0 人在线

发一些好玩的 现在成 mb 的私人频道了 Links t.me/Rosmontis_Daily t.me/PDChinaNews

成员规模13,581 位成员
在线情况0 人在线
消息总数1,042 条消息
浏览量总数1,360,923 次浏览

在这个频道里搜索消息……

t.me/microblock_pub

> We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications.
新活是一个支持图像多模态的 LLM,成功把图像生成和理解在单个模型中统一起来(不像其它大模型生成图片都用自然语言调用什么 SD Flux 啥的其它模型 ⁽¹⁾)
训练方式是传统 预训练 & SFT,没有用强化学习
这个模型比较小,只有 7b 参数量,大家可以随意本地运行,看这个 Series 估计先 PoC 以后后面再搞个大的
看技术报告里面全面打爆同参数量模型,技术报告还没上传,传了再看
现在预定的链接:
线上 Playground:https://huggingface.co/spaces/deepseek-ai/Janus-Pro-7B
技术报告:https://github.com/deepseek-ai/Janus/blob/main/janus_pro_tech_report.pdf
DeepSeek 到底在干嘛,除夕也有新活,这也卷?感觉可以给 DS 磕两个
再这样下去别人的新模型就要比不上baseline了
——————-
⁽¹⁾: Gemini 2.0/GPT4o 也有在实验中的类似原理的多模态,但还并未开放使用,此处并不意为所有大模型均为语言交互生图