Перейти к содержимому

Vectorized UDF: Scalable Analysis with Python and PySpark - Li Jin

Databricks

0:00 / 0:00

Vectorized UDF: Scalable Analysis with Python and PySpark - Li Jin

6 418 просмотров · 7 лет назад
Databricks
166 тыс. подписчиков
6 418 просмотров · 7 лет назад
Li Jin, a software engineer at Two Sigma shares a new type of Py Spark UDF: Vectorized UDF. Over the past few years, Python has become the default language for data scientists. Packages such as pandas, numpy, statsmodel, and scikit-learn have gained great adoption and become the mainstream toolkits. At the same time, Apache Spark has become the de facto standard in processing big data. Spark ships with a Python interface, aka PySpark, however, because Spark’s runtime is implemented on top of JVM, using PySpark with native Python library sometimes results in poor performance and usability. Vectorized UDF is built on top of Apache Arrow and bring you the best of both worlds – the ability to define easy to use, high performance UDFs and scale up your analysis with Spark. To learn more: https://databricks.com/blog/2017/10/3... About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unifie... Connect with us: Website: https://databricks.com Facebook:   / databricksinc   Twitter:   / databricks   LinkedIn:   / databricks   Instagram:   / databricksinc   Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-nam...