Multimodal research · Data infrastructure · Computer vision

James Hong

I build the data foundations behind image generation.

Research scientist at Reve AI, leading data infrastructure and curation for billion-scale multimodal datasets used in image generation and editing. PhD in Computer Science from Stanford University.

At Reve AI

Data systems for frontier image models

My work connects large-scale data infrastructure, model capabilities, and the creative needs of real products.

Infrastructure & curation

Billion-scale multimodal data

My team builds billion-scale multimodal datasets and infrastructure for pre-training state-of-the-art image generation and editing models. We curate high-quality data around aesthetics, style, control, quality, and safety, and fine-tune frontier-scale MLLMs for creative applications.

  • Multimodal infrastructure
  • Aesthetic curation
  • Product-driven datasets
  • MLLM fine-tuning

Prior research

Learning from unlabeled images and video

During my PhD, I designed systems and learning methods for understanding large, unstructured collections of images and video, with an emphasis on weak supervision.

Project preview for Learning Subject-Aware Cropping by Outpainting Professional Photos

AAAI Conference on Artificial Intelligence (AAAI) 2024

Learning Subject-Aware Cropping by Outpainting Professional Photos

James Hong, Lu Yuan, Michaël Gharbi, Matthew Fisher, and Kayvon Fatahalian

TL;DR: We turn a stock photo database into a weakly labeled dataset for learning what makes an aesthetically pleasing composition. Despite being only weakly supervised, our system outperforms supervised methods trained on large, crowd-annotated datasets.

Project preview for Spotting Temporally Precise, Fine-Grained Events in Video

European Conference on Computer Vision (ECCV) 2022

Spotting Temporally Precise, Fine-Grained Events in Video

James Hong, Haotian Zhang, Michaël Gharbi, Matthew Fisher, and Kayvon Fatahalian

TL;DR: We propose an efficient neural network for processing every frame of a video to detect very temporally fine-grained events at the granularity of a single frame.

Project preview for Analyzing the Faces in a Decade of US Cable TV News

Conference on Knowledge Discovery and Data Mining (KDD) 2021

Analyzing the Faces in a Decade of US Cable TV News

James Hong, Will Crichton, Haotian Zhang, Daniel Y. Fu, Jacob Ritchie, Jeremy Barenholtz, Ben Hannel, Xinwei Yao, Michaela Murray, Geraldine Moriba, Maneesh Agrawala, and Kayvon Fatahalian

TL;DR: We conduct an analysis of the visual content in 300,000 hours of US cable TV news. I built infrastructure to process the video and the TV news analyzer, a public interface and query engine for interactive search and visualization of all of the data.

Project preview for Learning in situ: A Randomized Experiment in Video Streaming

Symposium on Networked Systems Design and Implementation (NSDI) 2020

Learning in situ: A Randomized Experiment in Video Streaming

Francis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi, James Hong, Keyi Zhang, Philip Levis, and Keith Winstein

Awards: USENIX NSDI Community Award, IRTF Applied Networking Research Prize
TL;DR: A public research platform for conducting video streaming experimentation. I designed the JS/HTML web client, which streams video and audio chunks over a WebSocket connection.

More publications 6 projects in vision, systems, and IoT

SyntheticData4ML Workshop @ NeurIPS 2023

Learning to Place Objects into Scenes by Hallucinating Scenes around Objects

Lu Yuan, James Hong, Vishnu Sarukkai, and Kayvon Fatahalian

AI Systems Workshop @ SOSP 2019

Rekall: Specifying Video Events using Compositions of Spatiotemporal Labels

Dan Fu, Will Crichton, James Hong, Xinwei Yao, Haotian Zhang, Anh Truong, Avanika Narayan, Maneesh Agrawala, Christopher Ré, and Kayvon Fatahalian

Conference on Internet of Things Design and Implementation (IoTDI) 2018

Securing the Internet of Things With Default-Off Networking

James Hong, Amit Levy, Laurynas Riliskis, and Philip Levis

Conference on Internet of Things Design and Implementation (IoTDI) 2018

Tethys: Collecting Sensor Data Without Infrastructure or Trust

Holly Chiang, James Hong, Kevin Kiningham, Laurynas Riliskis, Philip Levis, and Mark Horowitz

Conference on Mobile Systems, Applications, and Services (MobiSys) 2016

Beetle: Flexible Communication for Bluetooth Low Energy

Amit Levy, James Hong, Laurynas Riliskis, Philip Levis, and Keith Winstein

IoT-App Workshop @ SenSys 2015

Ravel: Programming IoT Applications as Distributed Models, Views, and Controllers

Laurynas Riliskis, James Hong, and Philip Levis

Experience & education

From systems to multimodal research

2024 - present

Research Scientist at Reve AI

Head of Data (Curation and Infra)

2017 - 2024

Research Assistant at Stanford University

Computer Science, Graphics Lab

2020

Research Intern at Adobe

Creative Intelligence Lab

2016, 2017

Software Engineering Intern at Rubrik

Security team

2015

Software Engineering Intern at LinkedIn

Data Analytics Infrastructure

2014

Software Development Intern at PlayStation (SNEI)

Experimentation Platform

Travel and Hobbies

Photography

Photographs from travel, landscapes, architecture, and wildlife.