Deleting consecutive rows in a pandas dataframe with the same value

Question

How can I delete only the three consecutive rows in a pandas dataframe that have the same value (in the example below, this would be the integer "4").

Consider the following code:

import pandas as pd

df = pd.DataFrame({
    'rating': [4, 4, 3.5, 15, 5 ,4,4,4,4,4 ]
})

   rating
0  4.0
1  4.0
2  3.5
3  15.0
4  5.0
5  4.0
6  4.0
7  4.0
8  4.0
9  4.0

I would like to get the following result as output with the three consecutive rows containing the value "4" being removed:

ansev · Accepted Answer · 2022-02-22 12:34:52Z

2

first get a group each time a new value exists, then use GroupBy.head

new_df = df.groupby(df['rating'].ne(df['rating'].shift()).cumsum()).head(2)
print(new_df)

   rating
0     4.0
1     4.0
2     3.5
3    15.0
4     5.0
5     4.0
6     4.0

edited Feb 22, 2022 at 12:34

answered Feb 22, 2022 at 12:29

ansev

31k5 gold badges21 silver badges33 bronze badges

Sign up to request clarification or add additional context in comments.

Comments

jezrael · Accepted Answer · 2022-02-22 12:32:28Z

1

Use GroupBy.cumcount for counter and filter in rows in boolean indexing:

#filter consecutive groups less like 2 (python count from 0)
df= df[df.groupby(df['rating'].ne(df['rating'].shift()).cumsum()).cumcount().lt(2)]
print (df)
   rating
0     4.0
1     4.0
2     3.5
3    15.0
4     5.0
5     4.0
6     4.0

edited Feb 22, 2022 at 12:32

answered Feb 22, 2022 at 12:07

jezrael

868k103 gold badges1.4k silver badges1.3k bronze badges

Collectives™ on Stack Overflow

Deleting consecutive rows in a pandas dataframe with the same value

2 Answers 2

Comments

Comments

Your Answer

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

Comments

Comments

Your Answer

Sign up or log in

Post as a guest

Related