Extending Pandas using Apache Arrow and Numba - Uwe L Korn
PyData
0:00 / 0:00
Extending Pandas using Apache Arrow and Numba - Uwe L Korn
14 200 просмотров · 8 лет назад
PyData
173 тыс. подписчиков
14 200 просмотров · 8 лет назад
PyData Berlin 2018
With the latest release of Pandas the ability to extend it with custom dtypes was introduced. Using Apache Arrow as the in-memory storage and Numba for fast, vectorized computations on these memory regions, it is possible to extend Pandas in pure Python while achieving the same performance of the built-in types. In the talk we implement a native string type as an example.
Slides: https://pydata.org/berlin2018/proposa...
---
www.pydata.org
PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R.
PyData conferences aim to be accessible and community-driven, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases.
0:00 - Introduction
0:41 - Speaker's background
1:53 - Introduction to Pandas and NumPy
3:05 - Shortcomings of Pandas
8:55 - Extending Pandas with ExtensionArrays
12:16 - Apache Arrow for Data Storage
15:50 - Numba for computing
21:03 - Putting all together
28:10 - Closing
28:45 - Questions
S/o to https://github.com/keckelt for the video timestamps!
Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: https://github.com/numfocus/YouTubeVi...