Split set values from Pandas dataframe cell over multiple rows

Question

I have a pandas DataFrame in the following form:

    col1           col2
1    a       {hu, fdf, ko, dss}
2    b       {sdsjdn, lk}
3    c       {sds, aldj, dhva}

Now I want to split the set values over multiple rows to make it look like this:

    col1           col2
1    a              hu
2    a              fdf
3    a              ko
4    a              dss
5    b              sdsjdn
6    b              lk
7    c              sds
8    c              aldj
9    c              dhva

Anyone has any insights how I can do this?

jezrael · Accepted Answer · 2017-03-08 10:14:23Z

3

You need numpy.repeat for create new duplicated column with flattening of another set column by chain.from_iterable:

df = pd.DataFrame({ 'col1': ['a','b','c'],
                   'col2': [set({'hu', 'fdf', 'ko', 'dss'}),
                            set({'sdsjdn', 'lk'}),
                            set({'sds', 'aldj', 'dhva'})]})

print(df)
  col1                col2
0    a  {hu, dss, ko, fdf}
1    b        {lk, sdsjdn}
2    c   {dhva, aldj, sds}

from  itertools import chain

df1 = pd.DataFrame({
        "col1": np.repeat(df.col1.values, df.col2.str.len()),
        "col2": list(chain.from_iterable(df.col2))})

print (df1)
  col1    col2
0    a      hu
1    a     dss
2    a      ko
3    a     fdf
4    b      lk
5    b  sdsjdn
6    c    dhva
7    c    aldj
8    c     sds

answered Mar 8, 2017 at 10:14

jezrael

868k103 gold badges1.4k silver badges1.3k bronze badges

Sign up to request clarification or add additional context in comments.

Collectives™ on Stack Overflow

Split set values from Pandas dataframe cell over multiple rows

1 Answer 1

Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

1 Answer 1

Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related