Showing posts with label tr. Show all posts
Showing posts with label tr. Show all posts

Thursday, January 8, 2015

[Bash] Count Letters in a Text File

Bash script have very useful commands for Text File analysis. For this example we will use small part from Leo Tolstoy: War and Peace, available on Project Gutenberg:




This part is placed in file onepart.txt. To count letters all letters must have same case. Command tr will translate all letters to uppercase letters:
$ cat onepart.txt|tr a-z A-Z
Next, output is filtered. Command tr is useful here with first argument are switches -cd and seccond is letter to be filtered. Switch -c means complent and -d to delete, in short those two arguments means erase everything what is not letter in argument.
Pipelining we get:
$ cat onepart.txt|tr a-z A-Z|tr -cd A
AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 
Or for B, C etc...

$ cat 3.txt|tr a-z A-Z|tr -cd B
BBBB

$ cat 3.txt|tr a-z A-Z|tr -cd C
CCCCCCCCCCCCCCCCCCCCCCC

 $ cat 3.txt|tr a-z A-Z|tr -cd D
DDDDDDDDDDDDDDDDDDDDDDDDDD

And of course we need to count those letters. Command wc with switch -m do exactly that:

$ cat 3.txt|tr a-z A-Z|tr -cd A|wc -m
40

$ cat 3.txt|tr a-z A-Z|tr -cd B|wc -m
4

$ cat 3.txt|tr a-z A-Z|tr -cd C|wc -m
23

$ cat 3.txt|tr a-z A-Z|tr -cd D|wc -m
26
and, of course, script to count all letters occurances: 

#!/bin/bash

text=$(cat $1|tr a-z A-Z)
echo "Letter occurances:"

for l in A B C D E F G H I J K L M N O P Q R S T U V W X Y Z
do 
let=$(echo $text|tr -cd $l|wc -m)
echo "$l $let"
done
Script is a little bit slower because text is analyzed for every letter.

Sunday, October 14, 2012

Bash - Count words in text and make dictionary

In this tutorial i'll try to explain how to count words in some text and make list of used words.
First we could get some text file. I get War and Peace by Leo Tolstoy from Project Gutenberg:
wget http://gutenberg.org/files/2600/2600.txt

After getting file, we must to convert uppercase letters to lowercase:
tr A-Z a-z
then convert everything which is not small letter to new line
tr -cs a-z '\n'
then sort result
sort
finally get unique words and count them:
uniq -c
To get this into working condition we must put those commands into pipeline. To do that we'll make file wordcount:
nano wordcount
and type this:

cat "$@" | tr A-Z a-z | tr -cs a-z '\n'|sort|uniq -c
save and exit nano editor.
Give executable rights to file wordcount:
chmod +x wordcount
and count words with
./wordcount <textfile>
in this case:
./wordcount 2600.txt
and we'll get words used in this great book and their number of appearances.

If we want to get only dictionary used in this book, we make another file:
nano makedict
and type this:
cat "$@" | tr A-Z a-z | tr -cs a-z '\n'|sort|uniq
take note that only difference between wordcount and makedict scripts is -c switch in uniq command.
of course, we must give execution rights to file:
chmod +x makedict
and use:
./makedict 2600.txt