awk 依据几列的值删除重复行

依据几列的值删除重复行

输入
1.txt
a 1 1
a 1 2
a b 3
a c 4
a c 5
a c 6
a d 7
a d 8
a e 9
a f 10
a f 11

输出
2.txt
a b 3
a e 9

先把1.txt按照指定的几列(第1,2列)排序
sort -k1,1 -k2,2  1.txt > 3.txt

awk '{last=$0; chr=$1;start=$2; getline; if( chr == $1 && start == $2 ) {print last" 1";print $0" 1";} else {print last" 0";print $0" 0";}}' 3.txt
> 4.txt

awk 'BEGIN{getline;print $0" "0;}{last=$0; chr=$1;start=$2; getline; if( chr == $1 && start == $2 ) {print last" 1";print $0" 1";} else {print last" 0";print $0" 0";}}' 4.txt
> 5.txt

awk '$4 == 0 && $5 == 0'  5.txt | awk '{print $1" "$2" "$3;}' > 2.txt

原文地址:https://www.cnblogs.com/emanlee/p/15177785.html