pyspark program for nested loop

Question

I am new to PySpark and I am trying to understand how can we write multiple nested for loop in PySpark, rough high level example below. Any help will be appreciated.

for ( i=0;i<10;i++)
   for ( j=0;j<10;j++)
       for ( k=0;k<10;k++)
          { 
           print "i"."j"."k"
}

zero323 · Accepted Answer · 2015-07-20 19:19:05Z

8

In non distributed setting for-loops are rewritten using foreachcombinator, but due to Spark nature map and flatMap are a better choice:

from __future__ import print_function
a_loop = lambda x: ((x, y) for y in xrange(10))
print_me = lambda ((x, y), z): print("{0}.{1}.{2}".format(x, y, z)))

(sc.
    parallelize(xrange(10)).
    flatMap(a_loop).
    flatMap(a_loop).
    foreach(print_me)

Of using itertools.product:

from itertools import product
sc.parallelize(product(xrange(10), repeat=3)).foreach(print)

edited Jul 20, 2015 at 19:19

answered Jul 20, 2015 at 10:48

zero323

331k108 gold badges982 silver badges958 bronze badges

Sign up to request clarification or add additional context in comments.

1 Comment

10465355 Over a year ago

@Ali print in lambda work's just fine. Tuple parameter unpacking doesn't, but the OP explicitly uses Python 2.

Collectives™ on Stack Overflow

pyspark program for nested loop

1 Answer 1

1 Comment

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

1 Answer 1

1 Comment

Your Answer

Sign up or log in

Post as a guest

Linked

Related