跳到主要內容
Microsoft
separator
https://catalogartifact.azureedge.net/publicartifacts/kcloudhubllc1763357129530.pyspark1-9cba059e-b92a-4807-9700-a43791a8c85b/f44463af-9b26-4dfb-8d19-21d8db17870f_kcloudlogo.txt.png

Pyspark

作者 kCloudHub LLC

(1 評分)

Version 4.1.1 + Free Support on Ubuntu 24.04

PySpark is an open-source Python API for Apache Spark that enables large-scale data processing and distributed computing. It allows developers and data engineers to use Python to work with big data efficiently across clusters.

Key Features of PySpark:

  • Open-source distributed data processing framework with Python support.
  • Scalable processing of large datasets across clusters.
  • Built-in libraries for SQL queries, machine learning, and streaming.
  • High-performance in-memory data processing.
  • Integration with data sources such as Hadoop, cloud storage, and databases.

PySpark Usage:

$ sudo su
$ cd /opt/pyspark
$ source pyspark-env/bin/activate
$ python3 -c "import pyspark; print(pyspark.__version__)"
  

Disclaimer:
PySpark is an independent open-source project that is part of the Apache Spark ecosystem and is maintained by the Apache Software Foundation.

中文(台灣)
您的隱私權選擇退出圖示 您的隱私選擇
消費者健康情況隱私權 網站地圖 Contact Us 隱私權與 Cookie 使用條款 商標 關於我們的廣告 管理 Cookie