2018년 4월 7일 토요일

Hidden layer의 뉴런 개수 결정

Hidden layer의 뉴런 개수 결정

대부분의 문제는 히든 레이어가 1개면 충분하다.

히든 레이어의 개수가 다음에 해당하면 네트워크의 성능이 떨어진다.

  1. 뉴런이 1개인 경우
  2. 뉴런의 개수가 입력층의 뉴런 개수의 평균 또는 출력층의 뉴런 개수의 평균인 경우

뉴런의 개수 결정 방법에 대한 자세한 설명은 출처를 참고하라.

출처 : <https://stats.stackexchange.com/questions/181/how-to-choose-the-number-of-hidden-layers-and-nodes-in-a-feedforward-neural-netw>

2018년 3월 20일 화요일

딕셔너리 소팅하기

파이썬 딕셔너리 소팅하기

[질문]

  dict에서 키가 아닌 값으로 소팅하고 싶습니다 DB에서 2개 필드(string value, numeric value 하나씩)를 읽어 dictionary x = {'1': 2, '3': 4, '4': 3, '2': 1, '0': 0} 를 만들었습니다. string은 유니크한 키이고 numeric은 그렇지 않습니다. string 기준으로는 소팅할 수 있는데 numeric으로는 소팅하는 방법을 모르겠어요. 어떻게 해야되나요?


[답변]

   dict는 리스트나 튜플이랑은 다르게 순서가 없습니다. 따라서 dict 타입은 소팅할 수 없기때문에, dict를 소팅하고 싶다면 dict를 튜플의 리스트로 표현해야 합니다.
예를들어. 키가 아닌 값으로 소팅 하고 싶다면
import operator
x = {'1': 2, '3': 4, '4': 3, '2': 1, '0': 0}
sorted_x = sorted(x.items(), key=operator.itemgetter(1))
이 경우 sorted_x 는 2번째 요소를 기준으로 소팅된 튜플들의 리스트이고.
dict(sorted_x) 는 x랑 같습니다.
키로 소팅하고 싶다면 다음과 같이 쓸 수 있습니다.
import operator
x = {1: 2, 3: 4, 4: 3, 2: 1, 0: 0}
sorted_x = sorted(x.items(), key=operator.itemgetter(0))


출처: <http://hashcode.co.kr/questions/19/%ED%8C%8C%EC%9D%B4%EC%8D%AC-%EB%94%95%EC%85%94%EB%84%88%EB%A6%AC-%EC%86%8C%ED%8C%85%ED%95%98%EA%B8%B0>




dictionary의 items()와 iteritmes()의 차이점



dict.items() 와 dict.iteritems()는 파이썬 버전에 따라 다른 결과를 냅니다.
원래 파이썬 items()는 tuple을 원소로 가지는 list를 return 했습니다. 이 방법은 메모리가 많이 필요했기 때문에, generator가 도입된 후 메모리의 효율적 관리를 위해 items()대신 iterator-generator 메소드인 iteritems()를 쓰게 됩니다.

다만 2.x에서는 구버전과의 호환성을 위해서 items()와 iteritems()를 모두 지원했지만 python3에서는 iterm()은 list가 아닌 iterator를 return하고(python3 의 items() = python2의 iteritems()) iteritems() 메소드는 쓸 수 없습니다.

2018년 1월 18일 목요일

An unexpected error occurred on Navigator start-up psutil.AccessDenied (pid=725)

Anaconda Navigator Start Error

[검색 결과 #1]

에러 번호를 검색하여 찾은 것이나 정확하게 'pid=725' 항목은 아니었다. 하지만 따라하니 해결되었다.

please run

$ anaconda-navigator --reset

Did you try to run Navigator with sudo, and then without sudo (or Run as Administrator on Windows)?
Please update to the latest version of Navigator.
On Navigator click on the update button on the top right of the interface or
on the terminal type

$ conda update anaconda-navigator

출처 : <https://github.com/ContinuumIO/anaconda-issues/issues/1984>



[검색 결과 #2]

사실, 검색에서 'pid=725' 항목이 나온 결과도 있었으나, 따라해도 정상적으로 동작하지 않았다. 과정은 매우 유사한데 동작하지 않은 이유는 무엇일까?

Please remember to update to the latest version of Navigator to include
the latest fixes.

Open a terminal (on Linux or Mac) or the Anaconda Command Prompt (on windows)
and type:

$ conda update anaconda-navigator
$ conda update navigator-updater

출처 : <https://github.com/ContinuumIO/anaconda-issues/issues/7391>



2018년 1월 7일 일요일

2017년 12월 31일 일요일

Levinson-Durbin recursion

Matlab

levinson : Levinson-Durbin recursion

Syntax
a = levinson(r)
a = levinson(r, n)
[a, e] = levinson(r, n)
[a, e, k] = levinson(r, n)

Description
The Levinson-Durbin recursion is an algorithm for finding an all-pole IIR filter with a prescribed deterministic autocorrelation sequence. It has applications in filter design, coding, and spectral estimation. The filter that levinson produces is minimum phase.

a = levinson(r) finds the coefficients of a length(r)-1 order autoregressive linear process which has r as its autocorrelation sequence. r is a real or complex deterministic autocorrelation sequence. If r is a matrix, levinson finds the coefficients for each column of r and returns them in the rows of a. n=length(r)-1 is the default order of the denominator polynomial A(z); that is, a = [1 a(2) ... a(n+1)]. The filter coefficients are ordered in descending powers of z–1.

H(z) = 1/A(z)
A(z) = 1 + a(2)z^(−1) + ⋯ + a(n+1)z^(−n)

a = levinson(r,n) returns the coefficients for an autoregressive model of order n.
[a,e] = levinson(r,n) returns the prediction error, e, of order n.
[a,e,k] = levinson(r,n) returns the reflection coefficients k as a column vector of length n.

Example
>> data = [2, 2, 0, 0, -1, -1, 0, 0, 1, 1];
>> [r, lg] = xcorr(data, 'biased');
>> r(lg<0) = []
r =
1.2000  0.6000  0.0000  -0.3000  -0.6000  -0.3000  -0.0000  0.2000  0.4000  0.2000
>> [ar, e] = levinson(r, 3)
ar =
1.0000   -0.6250    0.2500    0.1250
e =
0.7875

출처 : <https://kr.mathworks.com/help/signal/ref/levinson.html>

Python

lazy_lpc Module : Linear Predictive Coding (LPC) module

github for AudioLazy : <https://github.com/danilobellini/audiolazy>
AudioLazy 0.6 Docs : <http://pythonhosted.org/audiolazy/index.html>

Module contents
ParCorError Error when trying to find the partial correlation coefficients (reflection
coefficients) and there’s no way to find them.
toeplitz Find the toeplitz matrix as a list of lists given its first line/column.
levinson_durbin Solve the Yule-Walker linear system of equations.
lpc This is a StrategyDict instance object called lpc. Strategies stored: 5.
parcor Find the partial correlation coefficients (PARCOR), or reflection
coefficients, relative to the lattice implementation of a given LTI FIR
LinearFilter with a constant denominator (i.e., LPC analysis filter, or
any filter without feedback).
parcor_stable Tests whether the given filter is stable or not by using the partial
correlation coefficients (reflection coefficients) of the given filter.
Find the Line Spectral Frequencies (LSF) from a given FIR filter.
lsf_stable Tests whether the given filter is stable or not by using the Line
Spectral Frequencies (LSF) of the given filter. Needs NumPy.

levinson_durbin(acdata, order=None)
acdata – Autocorrelation lag list, commonly the acorr function output.
order – The order of the resulting ZFilter object. Defaults to len(acdata) - 1.

Example
>>> data = [2, 2, 0, 0, -1, -1, 0, 0, 1, 1]
>>> acdata = acorr(data)
>>> acdata
[12, 6, 0, -3, -6, -3, 0, 2, 4, 2]
>>> ldfilt = levinson_durbin(acorr(data), 3)
>>> ldfilt
1 - 0.625 * z^-1 + 0.25 * z^-2 + 0.125 * z^-3
    >>> ldfilt.numerator
[1, - 0.625, 0.25, 0.125]
>>> ldfilt.error  # Squared! See lpc for more information about this
7.875


출처 : <http://pythonhosted.org/audiolazy/lazy_lpc.html#audiolazy.lazy_lpc.levinson_durbin>

2017년 12월 22일 금요일

파이썬 변수 명명 규칙 (PEP 8 -- Style Guide for Python Code)

PEP 8을 간단하게 요약하면

피해야 할 이름 :

  • 소문자 l, 대문자 O, 대문자 I하나만 변수의 이름으로 쓰는 것은 권장하지 않습니다. 특정 폰트에서 헷갈릴수도 있기 때문입니다.

패키지와 모듈의 이름 :

  • 모듈 이름은 짧아야 하고, 전부 소문자여야 합니다. 가독성을 위해서라면 밑줄(_)을 쓸 수 있습니다.
  • 패키지 이름 또한 짧아야 하고, 전부 소문자여야 합니다. 밑줄은 권장하지 않습니다

클래스 이름 :

  • 클래스 이름은 CapWords 형식(단어를 대문자로 시작)을 따릅니다

exception의 이름 :

  • exception은 클래스이므로, class와 동일하게 적용됩니다.
  • 다만, 맨 뒤는 "Error"로 끝나야 합니다.

전역변수의 이름 :
(전역 변수는 하나의 모듈 안에서만 쓰인다고 가정합니다)

  • 전역 변수의 이름을 짓는 것은 함수 이름을 짓는 것과 동일합니다.
  • from M import *과 같이 쓰일 모듈에서는 global이 export 될 것을 방지하기 위해 all`메커니즘이나 혹은 맨 앞을 밑줄로 시작해야 합니다.

함수의 이름 :

  • 함수의 이름은 원칙적으로 소문자여야 하고, 가독성을 위해서 밑줄(_)로 단어를 나눌 수 있습니다.
  • 간혹 threading.py같이 이미 대/소문자를 혼용하는 경우는 대/소문자를 같이 쓰는 경우도 있습니다.

함수와 메소드의 인자 :

  • 메소드 인스턴스에 쓰이는 첫 번째 인자는 무조건 self여야 합니다.
  • 클래스 메소드의 첫 번째 인자는 무조건 cls여야 합니다
  • 예약된 키워드(in 등)와 함수의 인자가 겹치는 경우, 변수 이름 맨 뒤에 밑줄 하나를 붙이는 것으로 대체합니다.(ex, class_)

메소드 이름과 인스턴스의 이름 :

  • 함수 이름과 동일합니다.
  • public이 아닌 메소드나 인스턴스의 이름은 밑줄로 시작합니다

상수의 이름 :

  • 상수 이름은 전부 대문자와 밑줄로 쓰는 것을 원칙으로 합니다.


출처: <http://hashcode.co.kr/questions/489/%ED%8C%8C%EC%9D%B4%EC%8D%AC%EC%97%90%EC%84%9C-%EB%B3%80%EC%88%98%ED%95%A8%EC%88%98-%EC%9D%B4%EB%A6%84%EC%9D%84-%EC%A7%80%EC%9D%84-%EB%95%8C-%EA%B7%9C%EC%B9%99%EC%9D%B4-%EC%9E%88%EB%82%98%EC%9A%94>

2017년 12월 20일 수요일

필터뱅크와 MFCC

Speech Processing for Machine Learning: Filter banks, Mel-Frequency Cepstral Coefficients (MFCCs) and What's In-Between

출처: <http://haythamfayek.com/2016/04/21/speech-processing-for-machine-learning.html> 


이 글에서는 필터 뱅크와 MFCC에 대해 논의하고 필터 뱅크가 인기를 얻는 이유에 대해 설명한다.
필터뱅크를 계산하는 것과 MFCC를 계산하는 것은 어느 정도 같은 과정을 포함한다. 두 경우 모두 필터뱅크는 계산하여야 하고, 약간의 추가적인 스탭을 거쳐 MFCC를 얻을 수 있다.

간단히 말하면,

  1. pre-emphasis filtering: s1 = pre-emphasis filter(s)
  2. Framing and windowing: s12 = framing(s1), s2 = windowing(s21)
  3. Fourier transform and Power spectrum: s3 = fft(s2), s4 = p_spec(s3)
  4. Computation of the filterbank: s5 = cal_filterbank(s4)
  5. Discrete Cosine Transform: s6 = dct(s5)
  6. mean normalization:
         Filterbank = mean_norm(s5)
         MFCC = mean_norm(s6)


1. pre-emphasis filtering

고주파를 증폭킴으로써 얻을 수 있는 장점:
  1. 고주파 성분은 저주파 성분에 비해 크기가 작기 때문에 이를 통해 주파수 스팩트럼의 밸러스를 맞출 수 있다.
  2. 푸리에 변환 중 발생할 수 있는 수치 문제를 피할 수 있다.
  3. SNR을 개선할 수 있다.
필터 식 : y(t)=x(t)− αx(t−1), 일반적으로 α는 0.95 또는 0.97
파이썬 코드 : 
   pre_emphasis = 0.97
   emphasized_signal = numpy.append(signal[0], signal[1:] - pre_emphasis * signal[:-1])


2. Framing and Windowing
음성 신호는 nonstationary하므로 이를 한꺼번에 푸리에 변환하는 것은 의미가 없다. 일반적으로 음성 처리에서,
  1. 프레임 길이 : 20ms~40ms 
  2. 오버랩 : 50% (+/- 10%)
  3. 해밍윈도우
파이썬 코드:
   frame_size = 0.025   # 25ms
   frame_stride = 0.01  # 10ms hop-size == 15ms overlap
   # Convert from seconds to samples.
   frame_length, frame_step = frame_size * sample_rate, frame_stride * sample_rate  
   signal_length = len(emphasized_signal)
   frame_length = int(round(frame_length))
   frame_step = int(round(frame_step))
   # Make sure that we have at least 1 frame.
   num_frames = int(numpy.ceil(float(numpy.abs(signal_length - frame_length)) / 
                       frame_step))  

   pad_signal_length = num_frames * frame_step + frame_length
   z = numpy.zeros((pad_signal_length - signal_length))
   # Pad Signal to make sure that all frames have equal number of samples without 
   # truncating any samples from the original signal.
   pad_signal = numpy.append(emphasized_signal, z) 
   indices = numpy.tile(numpy.arange(0, frame_length), (num_frames, 1)) +             
                numpy.tile(numpy.arange(0, num_frames * frame_step, frame_step), 
                (frame_length, 1)).T
   frames = pad_signal[indices.astype(numpy.int32, copy=False)]
   # Hamming window
   frames *= numpy.hamming(frame_length)
   # Explicit Implementation **
   # frames *= 0.54 - 0.46 * numpy.cos((2 * numpy.pi * n) / (frame_length - 1))  


3. Fourier transform and Power spectrum
N-point FFT을 적용. 일반적으로 N = 256 또는 512를 사용
Power spectrum (periodogram) = |FFT(x_i)|^2 / N

파이썬 코드:
   NFFT = 512
   mag_frames = numpy.absolute(numpy.fft.rfft(frames, NFFT))  # Magnitude of the FFT
   pow_frames = ((1.0 / NFFT) * ((mag_frames) ** 2))               # Power Spectrum


4. Filterbanks
Mel 스케일의 Trianglar filter (일반적으로 40 filters)를 Power spectrum에 적용

Filterbanks on the Mel-scale
Mel과 Hz의 관계식:
   f = 700(10^(m∕2595)−1)
   m = 2595 log_10⁡(1+f/700)  

Filterbank 식:
Filterbanks equation
파이썬 코드:
   low_freq_mel = 0
   nfilt = 22
   # Convert Hz to Mel
   high_freq_mel = (2595 * numpy.log10(1 + (sample_rate / 2) / 700))  
   # Equally spaced in Mel scale
   mel_points = numpy.linspace(low_freq_mel, high_freq_mel, nfilt + 2)  
   # Convert Mel to Hz
   hz_points = (700 * (10**(mel_points / 2595) - 1))  
   bin = numpy.floor((NFFT + 1) * hz_points / sample_rate)
   fbank = numpy.zeros((nfilt, int(numpy.floor(NFFT / 2 + 1))))
   for m in range(1, nfilt + 1):
      f_m_minus = int(bin[m - 1])   # left
      f_m = int(bin[m])                # center
      f_m_plus = int(bin[m + 1])    # right
      for k in range(f_m_minus, f_m):
         fbank[m - 1, k] = (k - bin[m - 1]) / (bin[m] - bin[m - 1])
      for k in range(f_m, f_m_plus):
         fbank[m - 1, k] = (bin[m + 1] - k) / (bin[m + 1] - bin[m])

   filter_banks = numpy.dot(pow_frames, fbank.T)
   # Numerical Stability
   filter_banks = numpy.where(filter_banks == 0, numpy.finfo(float).eps, filter_banks)  
   filter_banks = 20 * numpy.log10(filter_banks)  # dB


5. Mel-frequency Cepstral Coefficients (MFCCs)
이전 단계에서 계산한 필터 뱅크 계수는 상관도가 매우 높기때문에 머신러닝 알고리듬에서 문제가 될 수도 있다고 알려져있다. 따라서 필터 뱅크 계수의 상관도를 줄이기위해 DCT를 적용한다. 그러면 압축된 필터뱅크의 형태를 얻을 수 있다. 일반적으로 음성 인식에서는 2-13개만 남기고 나머지는 버린다.  이유는 나머지의 계수가 음성 인식에 크게 기여하지 못하기 때문이다. 
                                                                  
파이썬 코드:
   num_ceps = 12
   # Keep 2-13
   mfcc = dct(filter_banks, type=2, axis=1, norm='ortho')[:, 1 : (num_ceps + 1)] 
   (nframes, ncoeff) = mfcc.shape
   n = numpy.arange(ncoeff)
   lift = 1 + (cep_lifter / 2) * numpy.sin(numpy.pi * n / cep_lifter)
   mfcc *= lift  #*

6. Mean Normalization
스팩트럼의 밸런스를 유지하고 SNR을 향상시키기 위한 것으로, 모든 프레임에서 각 계수의 평균을 빼기만하면 된다.
                                                               
파이썬 코드:
   filter_banks -= (numpy.mean(filter_banks, axis=0) + 1e-8)
   mfcc -= (numpy.mean(mfcc, axis=0) + 1e-8)


Filterbanks vs. MFCC

필터뱅크를 계산하는 모든 과정은 음성신호와 인간의 인지에 대한 본성을 동기로 삼는다.
MFCC를 위한 추가적인 계산 과정은 머신러닝 알고리즘의 한계가 동기가 되었다.
MFCC에서 DCT는 필터뱅크 계수간의 상관도를 낮추는 과정으로 화이트닝(whitening)으로 간주된다. GMM-HMM이 유행할 때 MFCC도 유행하였다.
특히 MFCC와 GMMs-HMMs이 공동으로 자동음성인식의 표준기법 발전한 경우 매우 인기가 있었다.
음성인식 시스템에 딥러닝을 적용하는 요즘에는 상관도가 높은 입력에 덜 민감한 딥뉴럴네트워크에 과연 MFCC가 옳은 선택인가를 생각해볼 필요가 있다. 선형 변환인 DCT로 인해 비선형성이 강한 음성의 일부 정보가 사라지는 것은 바라는 바가 아니다.

푸리에 변환도 과연 필요한가?라는 질문도 할 수 있다. 푸리에 변환도 선형 변환이기 때문에 이를 무시하고 시간 영역의 신호를 직접 학습하는 것도 좋을 것이다. 사실 이미 이에 대한 긍정적인 연구 결과가 발표되었다.


결론

머신러닝 알고리즘이 상관도가 높은 입력에 영향을 받지 않는다면 Mel-scaled 필터뱅크를 사용하고,
머신러닝 알고리즘이 상관도가 높은 입력에 영향을 받기 쉬운 환경이라면 MFCC를 사용하라.


람다 표현식 (Lambda expression)

람다 표현식(Lambda expression)  람다 표현식으로 함수를 정의하고, 이를 변수에 할당하여 변수를 함수처럼 사용한다. (1) 람다 표현식       lambda <매개변수> : 수식      ※ 람다식을 실행하...