Falcon: Interactive Visual Analysis for Big Data

Last update: Dec 27, 2022

Related tags

Overview

Falcon: Interactive Visual Analysis for Big Data

Crossfilter millions of records without latencies. This project is work in progress and not documented yet. Please get in touch if you have questions.

The largest experiments we have done so far is 10M flights in the browser and ~180M flights or ~1.7B stars when connected to OmniSciDB (formerly known as MapD).

We have written a paper about the research behind Falcon. Please cite us if you use Falcon in a publication.

@inproceedings{moritz2019falcon,
  doi = {10.1145/3290605},
  year  = {2019},
  publisher = {{ACM} Press},
  author = {Dominik Moritz and Bill Howe and Jeffrey Heer},
  title = {Falcon: Balancing Interactive Latency and Resolution Sensitivity for Scalable Linked Visualizations},
  booktitle = {Proceedings of the 2019 {CHI} Conference on Human Factors in Computing Systems  - {CHI} {\textquotesingle}19}
}

Demos

1M flights in the browser: https://vega.github.io/falcon/flights/
7M flights in OmniSci Core: https://vega.github.io/falcon/flights-mapd/
500k weather records: https://vega.github.io/falcon/weather/

Usage

Install with yarn add falcon-vis. You can use two query engines. First ArrowDB reading data from Apache Arrow. This engine works completely in the browser and scales up to ten million rows. Second, MapDDB, which connects to OmniSci Core. The indexes are created as ndarrays. Check out the examples to see how to set up an app with your own data. More documentation will follow.

Features

Zoom

You can zoom histograms. Falcon automatically re-bins the data.

Show and hide unfiltered data

The original counts without filters, can be displayed behind the filtered counts to provide context. Hiding the unfiltered data shows the relative distribution of the data.

With unfiltered data.

Without unfiltered data.

Circles or Color Heatmap

Heatmap with circles (default). Can show the data without filters.

Heatmap with colored cells.

Vertical bar, horizontal bar, or text for counts

Horizontal bar.

Vertical bar.

Text only.

Timeline visualization

You can visualize the timeline of brush interactions in Falcon.

Falcon with 1.7 Billion Stars from the GAIA Dataset

The GAIA spacecraft measured the positions and distances of stars with unprecedented precision. It collected about 1.7 billion objects, mainly stars, but also planets, comets, asteroids and quasars among others. Below, we show the dataset loaded in Falcon (with OmniSci Core). There is also a video of me interacting with the dataset through Falcon.

Developers

Install the dependencies with yarn. Then run yarn start to start the flight demo with in memory data. Have a look at the other script commands in package.json.

Experiments

First version that turned out to be too complicated is at https://github.com/vega/falcon/tree/complex and the client-server version is at https://github.com/vega/falcon/tree/client-server.

Falcon: Interactive Visual Analysis for Big Data

Related tags

Overview

Falcon: Interactive Visual Analysis for Big Data

Demos

Usage

Features

Zoom

Show and hide unfiltered data

Circles or Color Heatmap

Vertical bar, horizontal bar, or text for counts

Timeline visualization

Falcon with 1.7 Billion Stars from the GAIA Dataset

Developers

Experiments

Owner

Vega

Lale is a Python library for semi-automated data science.

A crude Hy handle on Pandas library

PLStream: A Framework for Fast Polarity Labelling of Massive Data Streams

Spectacular AI SDK fuses data from cameras and IMU sensors and outputs an accurate 6-degree-of-freedom pose of a device.

PCAfold is an open-source Python library for generating, analyzing and improving low-dimensional manifolds obtained via Principal Component Analysis (PCA).

Nobel Data Analysis

Evidence enables analysts to deliver a polished business intelligence system using SQL and markdown.

Big Data & Cloud Computing for Oceanography

An interactive grid for sorting, filtering, and editing DataFrames in Jupyter notebooks

Bamboolib - a GUI for pandas DataFrames

Codes for the collection and predictive processing of bitcoin from the API of coinmarketcap

A data parser for the internal syncing data format used by Fog of World.

In this project, ETL pipeline is build on data warehouse hosted on AWS Redshift.

Intake is a lightweight package for finding, investigating, loading and disseminating data.

Kats, a kit to analyze time series data, a lightweight, easy-to-use, generalizable, and extendable framework to perform time series analysis, from understanding the key statistics and characteristics, detecting change points and anomalies, to forecasting future trends.

Python for Data Analysis, 2nd Edition

MDAnalysis is a Python library to analyze molecular dynamics simulations.

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

PySpark Structured Streaming ROS Kafka ApacheSpark Cassandra

DefAP is a program developed to facilitate the exploration of a material's defect chemistry