Member of Technical Staff - VLM
Salary
R199 688 – R1 970 803 per month
Job Type
full time
Posted
about 2 hours ago
Closing date
13 Oct 2026
Job Description
About the Opportunity
Black Forest Labs, the innovative minds behind groundbreaking technologies like Latent Diffusion, Stable Diffusion, and FLUX, are seeking a skilled and forward-thinking Member of Technical Staff to join their remote team. This role is central to advancing the state-of-the-art in generative AI, specifically focusing on multimodal vision-language models (VLMs). If you are passionate about pushing the boundaries of what's possible in AI-driven creativity and have a proven track record of significant contributions, this opportunity offers a chance to shape the future of image and video generation.
The company prides itself on a foundation of research excellence, an open science ethos, and a commitment to expanding human creativity. Their models are already empowering millions of creators, developers, and businesses globally, with FLUX representing one of the most sophisticated generative systems available today. This position is an invitation to become a key player in an environment that thrives on innovation and impactful scientific discovery.
Key Responsibilities
As a Member of Technical Staff, you will be instrumental in the development and refinement of advanced multimodal vision-language models. Your work will be deeply integrated into the FLUX technology stack, requiring you to not only implement but also innovate on novel architectures. A significant aspect of your role will involve designing sophisticated fine-tuning strategies. These strategies will be tailored to adapt general-purpose VLMs for specialised creative applications that go beyond standard capabilities, such as generating precise captions, interpreting editing instructions, or enhancing prompts in nuanced ways.
Furthermore, you will explore and implement integrations between VLM and Large Language Model (LLM) capabilities and the existing diffusion and flow pipelines. This cross-pollination of technologies aims to enhance generation quality and controllability without introducing computational bottlenecks. A crucial part of your responsibility will be to stay abreast of emerging multimodal architectures, rigorously evaluating them and translating cutting-edge research into tangible, practical improvements for the company's technologies.
What They're Looking For
Black Forest Labs is seeking an individual with a substantial background in pre-training or significantly advancing a vision-language model that has seen successful deployment in a production system or has been publicly released. This isn't a role for those who have only performed light fine-tuning or adaptation. Candidates should be able to demonstrate a strong publication record or an unambiguous production track record that clearly shows they are at the forefront of multimodal architecture development.
A deep, intuitive understanding of the interplay between vision and language representations is essential. This includes a solid grasp of concepts such as tokenisation, cross-modal alignment and grounding, and the mechanics of cross-modal attention, as well as an awareness of their potential failure modes. Experience with distributed training across multiple nodes at scale is a prerequisite for this position.
The ideal candidate will be comfortable operating at the crucial boundary between research and production, possessing a keen interest in ensuring that their work not only reads well in academic contexts but also ships effectively and generalises robustly in real-world applications. While not strictly mandatory, experience with diffusion or flow-based generative models is considered a strong asset, particularly if you have explored how autoregressive and diffusion paradigms can be effectively combined.
---
In South Africa, roles like this, focusing on advanced AI research and development, are typically found within specialised tech companies, research institutions, and sometimes larger corporate innovation arms. While salaries can vary significantly based on experience and specific company structures, individuals with a proven track record in cutting-edge AI, particularly in generative models, can command competitive remuneration packages. The path for such professionals often involves continuous learning, contributing to open-source projects, and potentially moving into more senior research or lead engineering positions.
Requirements
- You've pretrained or significantly advanced a VLM (not just SFT'd or LoRA'd one) that was deployed in a production system or released publicly
- Strong publication record or unambiguous production track record showing you push the frontier on multimodal architectures
- Deep understanding of how vision and language representations interact: tokenization, alignment, grounding, cross-modal attention, and the failure modes of each
- Experience with distributed training at multi-node scale
- Comfortable at the research/production boundary — you care whether the work ships and generalizes, not just whether it reads well
- Experience with diffusion or flow-based generative models is a strong plus — especially if you've thought about how autoregressive and diffusion paradigms can compose
About the employer
blackforestlabs
blackforestlabs is a hiring organisation operating in Remote within the research sector. They are currently recruiting for the Member of Technical Staff - VLM role advertised on this page. Visit the official application link for more about the company, its culture and the team you would be joining.
Interested in this role at blackforestlabs?
JobVault never charges job seekers to apply.
More Research jobs
See all →Get ready for your application
Free career guides written for South African job seekers.
