Showing posts with label awk. Show all posts
Showing posts with label awk. Show all posts

Tuesday, April 16, 2013

Command line chaining sure is nice...

Every now and then I will need to validate code changes that will update a bunch of data files (200 or so) where each data file is fairly large (over 100 MB).  I could use vi/vim to open one of the files to do searching, but that takes quite a while.  Most of my text editors won't handle such large files.  It just so happens that the files are tab delimited, and therefore easy to parse and read with awk and grep.

If the data looks like this (but with millions of permutations of something similar across a couple hundred files):

datavalue1<TAB>datavalue2<TAB>SpecialFieldValue1<TAB>1<TAB>datavalue3
datavalue4<TAB>datavalue5<TAB>SpecialFieldValue2<TAB>2<TAB>datavalue6
datavalue7<TAB>datavalue8<TAB>SpecialFieldValue1<TAB>3<TAB>datavalue9

And I want to only see the values for the fourth column for all rows that have the SpecialFieldValue2, then I will use a command similar to this:

grep -P '\tSpecialFieldValue2\t' * | awk '{print $4}' > SpecialFieldValue2_values.txt

The -P tells grep to use Perl style regular expression, so I can use '\t' to represent tab characters. The * is the file name mask, so this will grep every file in the current directory.

I can then look through the SpecialFieldValue2_values.txt file to see that the data is what I expected.


Thursday, March 21, 2013

The return of Super Sed and Wonder Awk...

I needed to compare the tabbed separated data of a file (file A) to expected data (file B).  However, the contents of file A contained the processing date in each line of output in the file.

To handle the issue of non-matching dates, the "processing date" in file B was updated to be the string "PROCESSING_DATE".  That just left the date in file A to contend with.

Here is where sed and awk came to the rescue.  I used head -n1 to get the first line of file A, and used awk to get the processing date (which appeared in the 11th column).  The processing date was stored in a variable named target_date. Next, I used sed to do a replacement on all instances of target_date in file A.  After which I was able to do a diff on the two files to see if the output was as expected.

Here is how it looked in the shell script:

# get the target date
target_date=`head -n1 fileA.txt |  awk '{print $11}'`
# get the sed argument using the target date
sed_args="s/$target_date/PROCESSED_DATE/g"

# do an inline replacement on the target file
sed -i $sed_args fileA.txt

# check for differences
diffs=`diff fileA.txt fileB.txt`

if [ -n $diffs ]; then
    echo "There were differences between the files."
else
    echo "No differences were found."
fi