GraphX and Graph Processing Flashcards
7 cards from real Apache Spark practice questions. Tap to flip, then mark Knew It or Still Learning โ missed cards come back until you master them.
Read the first 7 GraphX and Graph Processing flashcards as text
What does the Triangle Count algorithm in GraphX compute for each vertex?
Answer: The number of triangles that the vertex participates in
GraphX's triangleCount() computes the number of triangles passing through each vertex, which is useful for measuring local clustering and network density.
What does `joinVertices` allow you to do in GraphX?
Answer: Update vertex attributes by joining graph vertices with values from an external RDD
joinVertices performs a left-outer join between the graph's VertexRDD and an external RDD[(VertexId, U)], applying a map function to update vertex attributes with the joined values.
What is the Label Propagation Algorithm (LPA) used for in GraphX?
Answer: Detecting communities or clusters in a graph
LPA in GraphX detects communities by having each vertex iteratively adopt the label most common among its neighbors until the labels stabilize.
How do you access the underlying distributed data of a GraphX graph as standard RDDs?
Answer: graph.vertices and graph.edges
A GraphX Graph exposes its data directly via `graph.vertices` (VertexRDD) and `graph.edges` (EdgeRDD), which can be used as regular Spark RDDs.
What is the key advantage of the `EdgePartition2D` strategy in GraphX over `EdgePartition1D`?
Answer: It reduces vertex replication by hashing both source and destination IDs when assigning edges to partitions
EdgePartition2D hashes both the src and dst vertex IDs together to assign edges, distributing vertex ghost copies more evenly and reducing total replication compared to 1D.
In Pregel's execution model as implemented in GraphX, what happens to a vertex that receives no messages in a superstep?
Answer: The vertex does not execute its vertex program and remains unchanged
Vertices that receive no messages in a superstep skip execution of their vertex program entirely, which is key to Pregel's efficiency by avoiding unnecessary computation.
What is the primary use-case limitation of GraphX compared to dedicated graph databases such as Neo4j?
Answer: GraphX is not designed for low-latency, transactional (OLTP-style) single-vertex lookups
GraphX excels at large-scale batch analytics (OLAP) over graphs but is not optimized for the low-latency, transactional queries that native graph databases handle natively.