Multimodal research · Data infrastructure · Computer vision
James Hong
Building the data foundations for image generation.
Member of Technical Staff at OpenAI. Formerly, Research Scientist at Reve AI, leading data infrastructure and curation for billion-scale multimodal datasets used in image generation and editing. PhD in Computer Science from Stanford University.
At Reve AI
Data systems for frontier image models
My work connects large-scale data infrastructure, model capabilities, and the creative needs of real products.
Billions-scale multimodal data at Reve
My team builds billion-scale multimodal datasets and infrastructure for pre-training state-of-the-art image generation and editing models. We curate high-quality data around aesthetics, style, control, quality, and safety, and fine-tune frontier-scale MLLMs for creative applications.
Experience & education
From systems to multimodal research
2017 - 2024
Research Assistant at Stanford University
Computer Science, Graphics Lab
2020
Research Intern at Adobe
Creative Intelligence Lab
2016, 2017
Software Engineering Intern at Rubrik
Security team
2015
Software Engineering Intern at LinkedIn
Data Analytics Infrastructure
2014
Software Development Intern at PlayStation (SNEI)
Experimentation Platform
Prior research
Learning from unlabeled images and video
During my PhD, I designed systems and learning methods for understanding large, unstructured collections of images and video, with an emphasis on weak supervision.
AAAI Conference on Artificial Intelligence (AAAI) 2024
Learning Subject-Aware Cropping by Outpainting Professional Photos
TL;DR: We turn a stock photo database into a weakly labeled dataset for learning what makes an aesthetically pleasing composition. Despite being only weakly supervised, our system outperforms supervised methods trained on large, crowd-annotated datasets.
European Conference on Computer Vision (ECCV) 2022
Spotting Temporally Precise, Fine-Grained Events in Video
TL;DR: We propose an efficient neural network for processing every frame of a video to detect very temporally fine-grained events at the granularity of a single frame.
International Conference on Computer Vision (ICCV) 2021
Video Pose Distillation for Few-Shot, Fine-Grained Sports Action Recognition
TL;DR: We generate large amounts of dense weak-supervision on unlabeled and untrimmed sports video in order to learn more robust features for recognizing fine-grained actions.
Conference on Knowledge Discovery and Data Mining (KDD) 2021
Analyzing the Faces in a Decade of US Cable TV News
TL;DR: We conduct an analysis of the visual content in 300,000 hours of US cable TV news. I built infrastructure to process the video and the TV news analyzer, a public interface and query engine for interactive search and visualization of all of the data.
Symposium on Networked Systems Design and Implementation (NSDI) 2020
Learning in situ: A Randomized Experiment in Video Streaming
Awards: USENIX NSDI Community Award, IRTF Applied Networking Research Prize
TL;DR: A public research platform for conducting video streaming experimentation. I designed the JS/HTML web client, which streams video and audio chunks over a WebSocket connection.
More publications 6 projects in vision, systems, and IoT
Conference on Internet of Things Design and Implementation (IoTDI) 2018
Securing the Internet of Things With Default-Off Networking
Conference on Internet of Things Design and Implementation (IoTDI) 2018
Tethys: Collecting Sensor Data Without Infrastructure or Trust
Conference on Mobile Systems, Applications, and Services (MobiSys) 2016
Beetle: Flexible Communication for Bluetooth Low Energy
IoT-App Workshop @ SenSys 2015
Ravel: Programming IoT Applications as Distributed Models, Views, and Controllers
Travel and Hobbies
Photography
Photographs from travel, landscapes, architecture, and wildlife.